I've been comparing Mimo 2.6 Flash against GLM 5.3 Flash all week. On the 2 DGX Spark clusters, Mimo is definitely a bit faster, and on some tasks it does a better job. For example, here's a little 1-off vibe coded 3D racing game:
https://com-pute.com/nick/3D_racing--mimo2.6f--turbo-circuit-3d.html
That's better than most of the frontier models could do at the beginning of the year.
For all the obscure knowledge questions I've tried with those 2 models, GLM3 Flash just seems to know more details, even at IQ3 quant - but for most situations, that simply doesn't matter much, because all these models are great at compiling research online. Even models from 6 months ago do great at looking up info, if there's an Internet connection.
So for the moment, I'm keeping 1 DGX cluster running GLM Flash and another running Mimo Flash, to see if there's any clear consensus. As it stands now, they both feel great.
BTW, I'm still eagerly awaiting version 4 of the current Qwen 3.8 Flash Next architecture to be released. It runs faster than both MiMo and GLM Flash, at 4 bit quant, on a single 128Gb Strix Halo, DGX Spark, Mac Studio, etc. Its file size and memory use is much smaller than Mimo or GLM Flash, but even the current 3.8 Flash Next preview model seems to have genuinely comparable world knowledge, coding skills, and general agentic capability. In a few cases, Qwen 3.8 Flash Next has downright beaten Mimo and GLM at challenging coding tasks. It's now the default model that I keep loaded on Strix Halo machines.
All these best of breed free local models are already world changing - it's so exciting to imagine what we'll be running on local hardware by this time next year.