Post History

Current version by Nick Antonaccio

Current VersionAug 31, 2026 at 13:33

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only consisting of 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and many published experiences of other users (there's no shortage of feedback - OxAlpha was the #1 most used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GLM 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware. The most popular hardware standard right now for local clustered inference seems to be 2 DGX Spark machines. I use 2 Asus GX10s, connected directly with a QSFP112 DAC cable. That setup can currently be purchased for total of around $8700, delivered from Amazon. My next clustered tests will be with 2 Strix Halo machines, which would likely produce similar results (in the same speed ballpark, but with slower prefill speed), and can currently be purchased all-in for about $6500.

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Strix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become any extremely popular go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously inexpensive for frontier quality intelligence).

I'll post future updates about tests and general outcomes with this model here.

Previous Versions
Version 9Aug 31, 2026 at 13:33

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only consisting of 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and many published experiences of other users (there's no shortage of feedback - OxAlpha was the #1 most used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GLM 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware. The most popular hardware standard right now for local clustered inference seems to be 2 DGX Spark machines. I use 2 Asus GX10s, connected directly with a QSFP112 DAC cable. That setup can currently be purchased for total of around $8700, delivered from Amazon. My next clustered tests will be with 2 Strix Halo machines, which would likely produce similar results (in the same speed ballpark, but with slower prefill speed), and can currently be purchased all-in for about $6500.

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Strix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become any extremely popular go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously cheap for frontier quality intelligence).

I'll post future updates about tests and general outcomes with this model here!

Version 8Aug 31, 2026 at 13:32

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only consisting of 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and many published experiences of other users (there's no shortage of feedback - OxAlpha was the #1 most used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GLM 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware. The most popular hardware standard right now for local clustered inference seems to be 2 DGX Spark machines. I use 2 Asus GX10s, connected directly with a QSFP112 DAC cable. That setup can currently be purchased for total of around $8700, delivered from Amazon. My next clustered tests will be with 2 Strix Halo machines, which would likely produce similar results (in the same speed ballpark, but with slower prefill speed), and can currently be purchased all-in for about $6500.

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Stix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become any extremely popular go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously cheap for frontier quality intelligence).

I'll post future updates about tests and general outcomes with this model here!

Version 7Aug 30, 2026 at 20:00

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and the many published experiences of other users (and there's no shortage of feedback - OxAlpha was the #1 top used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GL 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware (the standard right now is 2 DGX Sparks (I use 2 clustered Asus GX10s, which you can currently find for around $8000 total, delivered from Amazon).

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Stix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become the go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!! - that's outrageously cheap for frontier quality intelligence).

Version 6Aug 30, 2026 at 15:30

I'm starting to use local GLM 5.3 Flash more for some actual work - it will likely end up replacing Deepseek V4 Flash as my default local workhorse model for challenging tasks.

The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. I've never seen any model close to that small size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, genuinely comparable in capability to the 2+ trillion parameter models (Kimi K3, Qwen 3.8 max, and the commercial models from OpenAI, Anthropic, etc.). This feeling is certainly backed up by benchmark scores and the many published experiences of other users (and there's no shortage of feedback - OxAlpha was the #1 top used model on OpenRouter, while it was being previewed for free, so there are a huge number of reviews).

GLM 5.3 Flash really took the thunder away from Qwen3.8-Flash-Next, which feels much more like a not-so-well-trained preview model, in comparison. That outcome was truly surprising to me.

GLM 5.3 Flash is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, GL 5.3 Flash is the most powerful self-hostable model, which can be run on less than $10,000 of local hardware (the standard right now is 2 DGX Sparks (I use 2 clustered Asus GX10s, which you can currently find for around $8000 total, delivered from Amazon).

Qwen 3.8 27b, Qwen 3.6 35a3, Gemma 4, and other models such as Mimo 2.5 still have their place on smaller consumer GPUs, and even on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to daily driver for agentic work, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Stix Halo, and other 128Gb machines - but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac or other platforms with 128Gb RAM, GLM 5.3 Flash will likely become the go-to model for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using API providers, but that's mostly just because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (I'll abbreviate it 'glm53f') to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used model. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!!).

Version 5Aug 30, 2026 at 15:29

I'm starting to use local GLM 5.3 Flash more for some actual work - it may end up replacing Deepseek V4 Flash as my default local workhorse model for challenging work. The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. I've never seen any model close to that size provide so much correct information about obscure topics (and not just the common topics which are often covered in training corpora).

GLM 5.3 Flash really feels like a frontier LLM, comparable with the 2+ trillion parameter models, and this feeling is certainly backed up by benchmarks and the experiences of other users (and there's no shortage of feedback - OxAlpha was the #1 top used model on OpenRouter, while it was being previewed for free).

This model is also natively multimodal, whereas Deepseek V4 Flash is not, and DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact.

As it stands, this is the most powerful self-hostable model which can be run on less than $10,000 of local hardware. Qwen 3.8 27b, Qwen 3.6 35a3, and Gemma models still have their place on smaller consumer GPUs, and on 128Gb machines, because they run faster, and require fewer resources. Deepseek V4 Flash is still my go-to, but only because it's been trusted in practice more than any other model at this point, and because it can run on a single DGX Spark, Stix Halo, and other 128Gb machines, but if the community ends up getting GLM 5.3 Flash running reliably on a single DGX Spark, Strix Halo, Mac and other platforms with 128Gb RAM, this model will likely become the go-to for local inference.

I'm still running Deepseek V4 Flash as my default model for agentic work, when using the API providers, but that's mostly because it's fully trusted after months of use, and because Cline-pass hasn't yet added GLM 5.3 Flash (glm53f) to the list of models it serves. Once Cline-pass adds it, I expect glm53f will become my most used models. It's actually now much cheaper on OpenRouter than Deepseek V4 Flash ($0.05/M input tokens $0.1667/M output tokens!!!)

Version 4Aug 30, 2026 at 15:19

I'm starting to use local GLM 5.3 Flash more for some actual work - it may end up replacing Deepseek V4 Flash as my default local workhorse model for challenging work. The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. GLM 5.3 Flash really feels like a frontier LLM, comparable with the 2+ trillion parameter models, and this feeling is backed up by the benchmarks. It's also natively multimodal, whereas DeepSeek-V4-Flash-Vision-Exp had vision capabilities tacked on after the fact. If the community ends up getting this model to run on a single DGX Spark, Strix Halo, Mac and other platforms with 128Gb RAM, it will be unstoppable.

Version 3Aug 30, 2026 at 14:01

I'm starting to use local GLM 5.3 Flash more for some actual work - it may end up replacing Deepseek V4 Flash as my default local workhorse model for serious work. The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. GLM 5.3 Flash really feels like a frontier model, and this feeling is backed up by the benchmarks. It's also natively multimodal, which beats the vision version of Deepseek Flash, which had vision capabilities tacked on after the fact. If the community ends up getting this model to run on a single DGX Spark, Strix Halo, Mac and other platforms with 128Gb RAM, it will be unstoppable.

Version 2Aug 30, 2026 at 13:58

I'm starting to use local GLM 5.3 more for some actual work - it may end up replacing Deepseek V4 Flash as my default local workhorse model for serious work. The model is not only top tier at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. GLM 5.3 Flash really feels like a frontier model, and this feeling is backed up by the benchmarks. It's also natively multimodal, which beats the vision version of Deepseek Flash, which had vision capabilities tacked on after the fact. If the community ends up getting this model to run on a single DGX Spark, Strix Halo, Mac and other platforms with 128Gb RAM, it will be unstoppable.

Version 1Aug 30, 2026 at 13:56

I'm starting to use local GLM 5.3 more for some actual work - it may end up replacing Deepseek V4 Flash as my default local workhorse model for serious work. The model is not only world class at coding, its world knowledge is also ridiculously deep, despite only being 320 billion parameters. GLM 5.3 Flash really feels like a frontier model (and this feeling is backed up by the benchmarks). It's also natively multimodal, which beats the vision version of Deepseek Flash, which had vision capabilities tacked on after the fact. If the community ends up getting this model to run on a single DGX Spark, Strix Halo, Mac and other platforms with 128Gb RAM, it will be unstoppable.