GLM 5.3 FLASH (OxAlpha) is a leader in local models. It's inexpensive + smart + fast on the API, originally an industry changing release, and still a killer buy.

436 views
Nick Antonaccio
Nick AntonaccioAdmin
Sep 02, 2026 at 15:39 (edited, 4 revisions)
#1

It turns out that the outrageously popular free stealth model OxAlpha was actually https://openrouter.ai/z-ai/glm-5.3-flash

(see https://www.linkedin.com/posts/phildour_glm-53-flash-just-came-out-of-stealth-it-activity-7498468815823933441-Gse3)

This model beats Opus 4.8 on most benchmarks, and if you've seen anything about OxAlpha, it is extremely capable in most ways that matter. You can use it yourself right now in Pi for only $0.075/$0.25 per million tokens in/out. Paste this prompt into Pi to add it :

Please add to pi the z-ai/glm-5.3-flash model from the openrouter provider. Please set the context length to 1 million tokens and the output length to the max allowable

The only other near-frontier model which competes with it in price is meta/muse-spark-1.2-contributor ($0.10/$0.20 per million tokens in/out) - that is the same model as meta/muse-spark-1.2, but you must agree to let Meta train on any data you enter into the contributor version, and it requires signing an 18 years old or older agreement.

You can run GLM 5.3 Flash locally

The really exciting thing about GLM 5.3 Flash is that it will run on 2 clustered DGX Spark machines, at a 4 bit quant, with enough VRAM to spare for a large KV cache. And it should run quickly, at something like 40 tokens per second, using the built-in ConnectX-7 link configured with RoCE v2 (running in vLLM with Ray).

Stay tuned for my tests!

Nick Antonaccio
Nick AntonaccioAdmin
Sep 02, 2026 at 15:42 (edited, 5 revisions)
#2

This model really needs 2 clustered DGX sparks to run. The Unsloth IQ3_XXS quant operates in llama.cpp with the full 1 million context limit, and it's still amazingly reliable at IQ3 compression (see some output examples below).

I'm getting speeds in the ballpark of 20 tokens per second running this model on llama.cpp.

Hopefully we'll be able to run GLM 5.3 Flash in vLLM soon, for faster performance. IQ4_XS runs well in llama.cpp, with higher precision, but I wouldn't push that quant beyond 128k context.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 31, 2026 at 13:33 (edited, 9 revisions)
#3

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only consisting of 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and many published experiences of other users (there's no shortage of feedback - OxAlpha was the #1 most used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GLM 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware. The most popular hardware standard right now for local clustered inference seems to be 2 DGX Spark machines. I use 2 Asus GX10s, connected directly with a QSFP112 DAC cable. That setup can currently be purchased for total of around $8700, delivered from Amazon. My next clustered tests will be with 2 Strix Halo machines, which would likely produce similar results (in the same speed ballpark, but with slower prefill speed), and can currently be purchased all-in for about $6500.

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Strix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become any extremely popular go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously inexpensive for frontier quality intelligence).

I'll post future updates about tests and general outcomes with this model here.

Nick Antonaccio
Nick AntonaccioAdmin
Sep 09, 2026 at 16:09 (edited, 14 revisions)
#4

Here are some quick example code generations by the Unsloth IQ3_XXS quant of GLM 5.3 Flash. This is the only quantization I tested yet, which I can tell you reliably leaves enough VRAM free to handle the full 1 million token KV cache, in llama.cpp:

Here's an example output generation, for a few questions I use to test world knowledge:

To be very clear, these responses were produced entirely without any research online or any other tool use by the model. It's genuinely staggering that those results came entirely from the information contained in 320 billion parameters, running locally on two relatively affordable boxes, pulling just a couple hundred watts (no datacenter-class hardware or electricity draw required). You could run this model entirely off-grid on a small solar panel.

The sort of detail about the obscure topics above, such as info about paramotor manufacturers, specific equipment models, and people, the jam.py framework, and my Rebol tutorials, has only ever been matched/exceeded by Fable (a 10 trillion parameter model which costs $50 per million to run on the API, compared to 25 cents per million (Fable is 200x more expensive to use, and not locally hostable)).

The details about my old and obscure Rebol, RFO, and other programming tutorials, makes me realize just how astoundingly broad and deep the knowledge in this model is. All other local models utterly fail in comparison. Models like Qwen 27b, for example, know absolutely nothing about these topics, or others of equal obscurity. It's going to be fun exploring just how much of all human knowledge this little self-hosted model knows.

What really blows me away is that these generations were all from a model that should have been significantly hobbled by 3 bit compression. Not only did glm35f have the information, it retrieved all the details correctly, at a level of compression that typically makes models experience lots of errors.

And the Rubik's cube example code above was produced first-shot - that was the output produced from a single prompt. No errors needed to be fixed, no iterations needed to be run to improve functionality.

This model is monstrously knowledgeable and capable. It seems to be better in every way than any of the frontier models which existed at the beginning of this year, commercial or open source. It's release represents a genuinely situation changing milestone in the industry.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 31, 2026 at 14:07
#5

It's amazing to note that Colibri already supports running GLM 5.3 Flash on machines which don't have a GPU, with as little as 25 GB of system RAM.

This architecture is just a proof of concept at this point - it runs extremely slowly, averaging somewhere around .04 tokens per second on a cold start. After warming up cache, output can conceivably get as high as 1-5 tokens per second.

It's just exciting to see a truly frontier quality model run, however slowly, without a GPU. Since this model is MIT licensed, we can expect improvements to CPU-only inference solutions.

A 1.58-bit ternary version, for example, would shrink this model's footprint to 60-70Gb. With 4.8 GB of data from an SSD per token, a ternary version would only need to read roughly 1.0 to 1.3 GB per token, potentially speeding up CPU generation times from 24 seconds per token to 4-6 seconds per token on standard NVMe drives. The issue is that ternary quantization frameworks (like BitNet 1.58b recipes) are designed for dense Transformer matrices, and GLM 5.3 Flash's highly complex, custom hybrid attention stack isn't well suited, but the industry is already working on mixed precision solutions, and we're just experiencing the first few steps of this long journey.

The exciting thing to me is to see real improvement, and actual output results on cheap hardware. Right now tools like Colibri are only potentially useful for tasks in which a page of output is produced overnight. The fact that that's possible, already, is astounding to me, and I'm optimistic to see models with current frontier capability eventually running directly on phones we carry in our pocket (or glasses, or whatever other mobile form factors evolve...).

Nick Antonaccio
Nick AntonaccioAdmin
Sep 02, 2026 at 15:47 (edited, 1 revision)
#6

Here's a little guitar chord diagram app generation completed on the Ollama Cloud API:

https://com-pute.com/nick/guitar-chord-maker--glm53f-ollama-cloud.html

It took just a fraction of a percent of weekly usage, after a bunch of iteration, to add features and adjust trial & error specifications about my own preferred functionality, output, and styling.

Nick Antonaccio
Nick AntonaccioAdmin
Sep 02, 2026 at 16:04 (edited, 2 revisions)
#7

Results on https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index agree with everything I've experienced about GLM 5.3 Flash, and which I've seen others report. It's a fantastic model, which occupies a unique position for anyone interested in self-hosting.

For my needs, preferences, and workflow patterns, large context tends to be very important, so I'm choosing to use the IQ3 quant. Keeping both IQ3 and IQ4 version on the same machine, however, and switching between them, is certainly not a problem. I just choose to leave IQ3 loaded and running all the time as the default version, and I can switch manually if the more precise version is ever needed. From my experience, though, IQ3 is working much more reliably than I would have expected. I'm starting to trust it for critical work. I suspect that this model is just so capable, that even when debilitated by very low quantization compression, it's still smarter and more reliable than alternative models. I've been impressed by how effectively the model thinks and evaluates its own output before returning results.

So far, aside from a pile of example/toy vibe coded application tests, like the ones above, I've jumped right into using GLM 5.3 Flash IQ3_XXS to produce anonymized data sets, which I feed into workflows that require HIPAA compliant data management. For example, when building apps with non-compliant ChatGPT, when the model needs to see actual data examples to build an application, that's my standard routine.

I do that many times a week, because it keeps me from having to spend ridiculous amounts of money on HIPAA compliant API providers, for software development tasks. GLM53f IQ3 has been doing a fantastic job providing that critical service for me. I've also started using it to perform all sorts of common IT tasks, which it's been breezing through without any trouble.

Please login to post a reply.

© 2026 AI By Nick.