I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.
The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only consisting of 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).
GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and many published experiences of other users (there's no shortage of feedback - OxAlpha was the #1 most used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).
GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.
GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.
As it stands, GLM 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware. The most popular hardware standard right now for local clustered inference seems to be 2 DGX Spark machines. I use 2 Asus GX10s, connected directly with a QSFP112 DAC cable. That setup can currently be purchased for total of around $8700, delivered from Amazon. My next clustered tests will be with 2 Strix Halo machines, which would likely produce similar results (in the same speed ballpark, but with slower prefill speed), and can currently be purchased all-in for about $6500.
Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Strix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become any extremely popular go-to model for local inference.
I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously inexpensive for frontier quality intelligence).
I'll post future updates about tests and general outcomes with this model here.