This repo contains specialized MoE-quants for XiaomiMiMo/MiMo-V2.6-Pro-MOPD. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

The MXFP4 quant is the "full quality" version, as the model has MXFP4 experts.

The BPW quants were produced using ed's bpw-size PR. The FFNs are the primarily quantized feature, rest of the model remains in Q8_0 / Q6_K

Quant Size Mixture PPL 1-(Mean PPL(Q)/PPL(base)) KLD
MXFP4 537.99 GiB (4.52 BPW) BF16 / MXFP4 3.171130 ยฑ 0.015401 +0.0352% -0.000000 ยฑ 0.000000
BPW3.5 416.74 GiB (3.50 BPW) Q8_0 / varies 3.254415 ยฑ 0.015658 +2.6625% 0.133381 ยฑ 0.000783
BPW3.0 357.11 GiB (3.00 BPW) Q8_0 / varies 3.479328 ยฑ 0.017232 +9.7575% 0.189016 ยฑ 0.001052
BPW2.5 297.63 GiB (2.50 BPW) Q6_K / varies 3.790107 ยฑ 0.019332 +19.5612% 0.261747 ยฑ 0.001379
BPW2.0 238.22 GiB (2.00 BPW) Q6_K / varies 4.554363 ยฑ 0.024378 +43.6701% 0.414109 ยฑ 0.001965

kld_graph ppl_graph

Downloads last month
2,349
GGUF
Model size
1T params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AesSedai/MiMo-V2.6-Pro-MOPD-GGUF

Quantized
(7)
this model

Space using AesSedai/MiMo-V2.6-Pro-MOPD-GGUF 1