Post History

Current version by Nick Antonaccio

Current VersionSep 02, 2026 at 15:42

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still amazingly reliable at IQ3 compression (see some output examples below).

I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.

Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.

Previous Versions
Version 5Sep 02, 2026 at 15:42

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still amazingly reliable at IQ3 compression (see some output examples below).

I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.

Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context. For my needs, preferences, and workflow patterns, large context tends to be very important, so I'm choosing to use the IQ3 quant.

Version 4Sep 02, 2026 at 15:41

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still amazingly reliable at IQ3 compression (see some output examples below).

I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.

Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.

Version 3Sep 02, 2026 at 15:40

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still pretty darn reliable at IQ3 compression (see some output examples below).

I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.

Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.

Version 2Aug 31, 2026 at 13:30

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still pretty darn reliable at IQ3 compression. I'm getting something in the ballpark of 20 tokens per second running it on llama.cpp. Hopefully we'll be able to run it on vllm soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.

Version 1Aug 30, 2026 at 19:46

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates with full 1 million context limit, and it's still pretty darn reliable in IQ3 compression. I'm getting something in the ballpark of 20 tokens per second running it on llama.cpp. Hopefully we'll be able to run it on vllm soon, for faster performance. IQ4_XS runs well, with higher precision, but I wouldn't push that quant beyond 128k context.