Post History

Current version by Nick Antonaccio

Current VersionSep 09, 2026 at 16:09

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models utterly fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It seems to be better in every way than any of the frontier models which existed at the beginning of this year, commercial or open source. It's release represents a genuinely situation changing milestone in the industry.

Previous Versions
Version 14Sep 09, 2026 at 16:09

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models utterly fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It seems to be better in every way than any of the frontier models which existed at the beginning of this year, commercial or open source. It's release represents a genuinely situation changing milestone in the industry.

Version 13Aug 31, 2026 at 13:42

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models absolutely fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It seems to be better in every way than any of the frontier models which existed at the beginning of this year, commercial or open source. It's release represents a genuinely situation changing milestone in the industry.

Version 12Aug 31, 2026 at 13:39

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models absolutely fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It's release represents a genuinely situation changing milestone in the industry.

Version 11Aug 31, 2026 at 13:36

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models absolutely fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It's release represents a genuinely situation changing milestone.

Version 10Aug 31, 2026 at 01:37

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models absolutely fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable.

Version 9Aug 31, 2026 at 01:36

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash (this is the only one I've tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp):

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use. It's genuinely staggering that those results came entirely from the information contained in a 320 billion parameter model, running locally on two little relatively affordable boxes (no datacenter-class hardware).

That sort of detail about those obscure topics, such as paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (200x more expensive, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials makes me realize just how astoundingly broad and deep the knowledge in this model is - it's going to be fun exploring just how much of all human knowledge this little self-hosted model knows. All other local models absolutely fail in comparison (models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity).

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did it have the information, it retrieved it correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality. This model is monstrously capable.

Version 8Aug 31, 2026 at 01:30

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash (this is the only one I've tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp):

Here's an example output generation, for a few questions I use to test world knowledge:

That's just a small sample, but in my tests none of the other models which can reasonably be run locally have been able to answer questions about obscure information as well.

Version 7Aug 30, 2026 at 20:05

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash (this is the only one I've tested which reliably leaves enough VRAM free to handle a 1 million token KV cache, in llama.cpp):

Here's an example output generation, for a few questions I use to test world knowledge:

That's just a small sample, but in my tests none of the other models which can reasonably be run locally have been able to answer questions about obscure information as well.

Version 6Aug 30, 2026 at 15:33

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash (this is the only one I've tested which reliably leaves enough VRAM free to handle a 1 million token KV cache, in llama.cpp):

https://com-pute.com/nick/rubiks--glm53f-iq3xxs.html https://com-pute.com/nick/space-invaders--glm53f-gx10.html

Here's an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

That's just a small sample, but in my tests none of the other models which run locally have been able to answer questions about obscure information as well.

Version 5Aug 30, 2026 at 15:06

Here's a quick example code generation by the Unsloth IQ3_XSS quant of GLM 5.3 Flash (this is the only one I've tested which reliably leaves enough VRAM free to handle a 1 million token KV cache, in llama.cpp):

https://com-pute.com/nick/space-invaders--glm53f-gx10.html

Here's an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

That's just a small sample, but in my tests none of the other models which run locally have been able to answer questions about obscure information as well.

Version 4Aug 30, 2026 at 14:34

Here's a quick example code generation by the Unsloth IQ3_XSS quant (this is the one that I've tested to leave enough VRAM free for 1 million token KV cache):

https://com-pute.com/nick/space-invaders--glm53f-gx10.html

And an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

That's just a small sample, but in my tests none of the other models which run locally have been able to answer questions about obscure information as well.

Version 3Aug 30, 2026 at 14:32

Here's a quick example code generation by the Unsloth IQ3_XSS quant (this is the one that I've tested to leave enough VRAM free for 1 million token KV cache):

https://com-pute.com/nick/space-invaders--glm53f-gx10.html

And an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

None of the other models which run locally have been able to answer questions about obscure information as well, in my tests.

Version 2Aug 30, 2026 at 14:03

Here's a quick example code generation by the Unsloth IQ3_XSS quant (this is the one that I've tested to leave enough VRAM free for 1 million token KV cache):

https://com-pute.com/nick/space-invaders--glm53f-gx10.html

And an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

None of the other models which run locally have been able to answer questions about obscure information as well.

Version 1Aug 30, 2026 at 13:58

Here's a quick example code generation:

https://com-pute.com/nick/space-invaders--glm53f-gx10.html

And an example world knowledge output:

https://com-pute.com/nick/glm53f-iq3xss-world-knowledge--pi-session-2026-08-30T13-09-54-023Z_01a052ca-70e7-793a-8719-ea881d598151.html

None of the other models that run locally have been able to answer questions about obscure information as well.