Post History

Current version by Nick Antonaccio

Current VersionSep 11, 2026 at 15:05

The truth is, so far, I genuinely dislike Qwen 3.8 27b. I know that's an unpopular opinion, but the proof has been in the pudding, in my experiences working with it.

I've run into more 'Response was truncated before completion' messages with it than any other popular model. It's taken so much tuning to get working somewhat reliably, on every machine I own, even the DGX Sparks and the Strix Halos. And I'm always still waiting for when it will fail on a task. The opposite is true of all the MOE models. Qwen 3.6 35a3 just competently zips along on every one of my 9 GPU servers, even the very old machines with 16Gb VRAM (q4 quants run great on those old, still relatively inexpensive to buy used machines).

Qwen 3.8 27b is certainly the smartest model that can be run on 'inexpensive' consumer hardware with low VRAM, but I'd prefer in every case to work with a faster MOE. Qwen 3.6 35a3 has served far better for my needs than 3.8 27b ever has. Perhaps it makes sense to put the time into configuring and tuning 3.8 27b on your hardware, if your goal is get the best possible first-shot visual design & layout results, or if you have absolutely no software development skills and just need a model to do everything for you. But with a little skill, and some time spent iterating through issues (a routine which everyone at the moment seems to be hell-bent on eliminating), I've gotten better, more reliable results from other models which can run on the same hardware. Do we really need to completely remove human input from every single task? The LLMs I use remove so much of the labor and time involved in completing virtually every development task, that software and IT work no longer feel like work at all - even though I still fully expect to do lots of iterations on every project I start. It seems that people want to just say a few words and not have to even look at anything but a first response from an agent. I don't have a desire to work that way. I want to be involved in every process, and use LLMs to take care of the overwhelming majority of tedious, time consuming work. So many models are already passing that bar. whatever single capability you think any 1 particular model offers, is likely not a real requirement. More skill and understanding is the requirement, and if we want human input to continue to be valuable, I think that continuing to develop skill, should be a requirement for anyone using AI tools.

If you have more powerful hardware, go with Deepseek V4 Flash, or even better GLM 5.3 Flash, if you've got the firepower.

It's becoming clear that hardware with 200Gb+ VRAM is what's required to really replace current frontier subscriptions, and there's still quite a gap between self-hosted models and the best API offerings. 512Gb VRAM closes a lot of that gap, but yet, still not entirely, when compared to Fable, Astra, and whatever comes next.

So I still use my $20 ChatGPT routine for all my most demanding software development projects, and I still never pay more than $20 per month. I would reach rate limits very quickly on the API and with Codex, but still never, in the ChatGPT interface. If/when that subsidized use disappears, I'm quite happy with GLM 5.3 Flash IQ3_XXS/IQ4_Xs and/or Deepseek Antirez/DS4 Q4 on two DGX Sparks. That hardware still can be gotten for $10,000ish, which is affordable enough for small-medium size businesses which require totally private self-hosted inference.

Qwen 3.8 Flash Next is their really exciting contender. That architecture seems to be far more effective, but the trained model that came with it is just a preview. I expect the Qwen 4 series will provide a much more solid foundation - let's hope they give us something that runs on a single machine with 128Gb VRAM.

I'm also very eager to see what Poolside delivers after the $7B commitment they got from Nvidia. Laguna S2.1 already runs at over 30 tokens per second on a single Strix Halo machine, and it's currently got a 1 million token context limit (compared to Qwen 3.8 27b's 256,000 limit). With real money and support from Nvidia, I expect Poolside will give us another new baseline for single 128Gb machines which people and small businesses can afford to buy without too much financial pain. Keep your eye on their next releases.

The last quarter of 2026 should be very exciting for self-hosting!

Previous Versions
Version 2Sep 11, 2026 at 15:05

The truth is, so far, I genuinely dislike Qwen 3.8 27b. I know that's an unpopular opinion, but the proof has been in the pudding, in my experiences working with it.

I've run into more 'Response was truncated before completion' messages with it than any other popular model. It's taken so much tuning to get working somewhat reliably, on every machine I own, even the DGX Sparks and the Strix Halos. And I'm always still waiting for when it will fail on a task. The opposite is true of all the MOE models. Qwen 3.6 35a3 just competently zips along on every one of my 9 GPU servers, even the very old machines with 16Gb VRAM (q4 quants run great on those old, still relatively inexpensive to buy used machines).

Qwen 3.8 27b is certainly the smartest model that can be run on 'inexpensive' consumer hardware with low VRAM, but I'd prefer in every case to work with a faster MOE. Qwen 3.6 35a3 has served far better for my needs that 3.8 27b ever has. Perhaps it makes sense to put the time into configuring and tuning 3.8 27b on your hardware, if your goal is get the best possible first-shot visual design and layout results, or if you have absolutely no software development skills and just need a model to do everything for you. But with a little skill, and some time spent iterating through issues (which everyone at the moment seems to be hell-bent on eliminating), I've gotten better, more reliable results from other models which can run on the same hardware. Do we really need to completely remove human input from every single task? The LLMs I use remove so much of the labor and time involved in completing virtually every task, that software development and IT work no longer feel like work at all, even though I still expect to do lots of iterations on every project I start. It seems that people want to just say a few words and not have to even look at anything but a first response from an agent. I don't even have a desire to work that way. I want to be involved in the process, and use LLMs to take care of the overwhelming majority of tedious, time consuming work. So many models are already passing that bar. You likely don't need whatever single capability you think any 1 particular model offers.

If you have more powerful hardware, go with Deepseek V4 Flash, or even better GLM 5.3 Flash, if you've got the firepower.

It's becoming clear that hardware with 200Gb+ VRAM is what's required to really replace current frontier subscriptions, and there's still quite a gap between self-hosted models and the best API offerings. 512Gb VRAM closes a lot of that gap, but still not entirely when compared to Fable, Astra, and whatever comes next.

So I still use my $20 ChatGPT routine for all my most demanding software development projects, and I still never pay more than $20 per month. I would reach rate limits very quickly on the API and with Codex, but still never, in the ChatGPT interface. If/when that subsidized use disappears, I'm quite happy with GLM 5.3 Flash IQ3_XXS on two DGX Sparks. That hardware still can be gotten for $10,000ish, which is affordable enough for small-medium size businesses which require totally private self-hosted inference.

Qwen 3.8 Flash Next is the really exciting contender. That architecture seems to be far more effective, but the trained model that came with it is just a preview. I expect the Qwen 4 series will provide a much more solid foundation - let's hope they give us something that runs on a single machine with 128Gb VRAM.

I'm also very eager to see what Poolside delivers after the $7B commitment they got from Nvidia. Laguna S2.1 already runs at over 30 tokens per second on a single Strix Halo machine, and it's currently got a 1 million token context limit (compared to Qwen 3.8 27b's 256,000 limit). With real money and support from Nvidia, I expect Poolside will give us another new baseline for single 128B-ish machines which people and small businesses can afford to buy without too much trouble. Keep your eye on their next releases.

The last quarter of 2026 should be very exciting for self-hosting!

Version 1Sep 11, 2026 at 14:56

The truth is, I dislike Qwen 3.8 27b. I know that's an unpopular opinion, but the proof has been in the pudding, in my experiences working with it.

I've run into more 'Response was truncated before completion' messages with it than any other popular model. It's taken so much tuning to get working somewhat reliably, on every machine I own, even the DGX Sparks and the Strix Halos. And I'm always still waiting for when it will fail on a task. The opposite is true of all the MOE models. Qwen 3.6 35a3 just competently zips along on every one of my 9 GPU servers, even the very old machines.

Qwen 3.8 27b is certainly the smartest model that can be run on 'inexpensive' consumer hardware, but I'd prefer in every case to work with a faster MOE. Qwen 3.6 35a3 has served far better for my needs that 3.8 27b ever has. Perhaps it makes sense to put the time in configuring and tuning 3.8 27b on your hardware, if the best possible first-shot visual design/layout results are what you care about, or if you have absolutely no software development skills and just need a model to do everything for you. But with a little skill and some human effort (which everyone seems to want to eliminate entirely at the moment), I've gotten better, more reliable results from other models which can run on the same hardware.

If you have more powerful hardware, go with Deepseek V4 Flash, or even better GLM 5.3 Flash, if you've got the firepower.

It's becoming clear that hardware with 200Gb+ VRAM is what's required to really replace frontier subscriptions, and there's still quite a gap. 512Gb VRAM closes a lot of that gap, but still not entirely when compared to Fable, Astra, and whatever comes next. So I still use my $20 ChatGPT routine for all my most demanding software development projects, and I still never pay more than $20 per month. If/when that subsidized use disappears, I'm quite happy with GLM 5.3 Flash IQ3_XXS on two DGX Sparks. That hardware still can be gotten for $10,000ish, which is affordable enough for small-medium size businesses which require self-hosted inference.

Qwen 3.8 Flash Next is the really exciting contender. That architecture seems to be far more effective, but the trained model that came with it is just a preview. I expect the Qwen 4 series will provide a much more solid foundation - let's hope they give us something that runs on a single machine with 128Gb VRAM.

I'm also very eager to see what Poolside delivers after the $5B they got from Nvidia. Laguna S2.1 already runs at over 30 tokens per second on a single Strix Halo machine, and it's got 1 million token context limit (compared to Qwen 3.8 27b's 256,000 limit). With real money and support from Nvidia, I expect they'll give us another new baseline for single machines that people can actually afford to buy. Keep your eye on their next releases. The last quarter of 2026 should be very exciting for self-hosting!