This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still amazingly reliable at IQ3 compression (see some output examples below).
I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.
Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.