Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:
- https://com-pute.com/nick/rubiks--glm53f-iq3xxs.html
- https://com-pute.com/nick/space-invaders--glm53f-gx10.html
- https://com-pute.com/nick/flying-game--glm35f.html
Here's an example output generation, for a few questions I use to test world knowledge:
To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.
The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).
The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models utterly fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.
What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.
And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.
This model is monstrously knowledgeable and capable. It seems to be better in every way than any of the frontier models which existed at the beginning of this year, commercial or open source. It's release represents a genuinely situation changing milestone in the industry.