I'm currently running Mimo 2.6 Flash on 2 clustered DGX Spark machines, at native shipped precision, ~20 tokens per second.
This is exciting because nothing was re-quantized, so this model should be as precise as it is on the API. That hasn't been the case for either Deepseek V4 Flash or GLM 5.3 Flash (both those models need to run at significant compression to enable speed and large context), so this model could potentially beat the quality of both those previous leaders in my locally hosted LLM stable - at the same speed.
The current setup is 4‑bit MXFP4 experts + FP8 E4M3 attention/dense + BF16 o_proj, FP8 KV cache. vLLM detects it as quantization=fp8 and runs the experts through its DEEPGEMM_MXFP4 MoE backend. No dequantization/GGUF step was involved.
I'll provide more feedback and examples soon!