Qwen 3.8 27b Issues

425 views
Nick Antonaccio
Nick AntonaccioAdmin
Aug 27, 2026 at 23:00 (edited, 9 revisions)
#1

Note from the future: see the recommended sampling settings in the post below - they make a difference. Initial issues appear to be resolving now. If you're using LM Studio, use the Unsloth versions, before the LM Studio Community version. Also be sure to try the Ridge quant by Empero. It runs on machines with as little as 12Gb VRAM.

Be sure to see all the info about the previous 3.6 versions of Qwen, plus Gemma 4: https://aibynick.com/thread/22

When you watch all the first review videos, everyone's response to Qwen 3.8 27b is as expected - it's being called the best model that can run on small consumer GPUs. And the benchmark hype all seems to show that it produces better output than 3.6, in most tests.

The only problem is that Qwen 3.8 27b thinks a lot. In fact, I'm experiencing significant issues with it terminating before completing tasks, because tasks are running so long. Perhaps another harness may help with that issue (the problematic terminations all occurred in Pi), but I've seen many reviews expressing a similar problem - even small software generations end up using nearly all the 256k context. Results are great, when the job is complete, but I'm regularly seeing jobs not completing.

To be clear, I'm not experiencing long context termination issues with only machines that have small GPUs. The task termination issue has occurred on both DGX Spark and Strix Halo machines, using both the q4 and q6 quants of 3.8 27b. So, this is not an isolated issue with just one version, one machine, or one configuration.

The other issue is that the 27b model is slow for a smallish model. The q4 quant on Strix Halo and the q6 quant on DGX Spark both run at just over 12 tokens per second. That's usable, but no where near the 50-60 tps I typically get with the Qwen 3.6 35a3 MOE model, on those same hardware platforms (and on faster hardware such as RTX 5090 and RTX 6000, the 3.6 MOE model runs over 200 tps).

I understand that historically, the dense 3.6 27b (dense) model is generally expected to produce higher quality output than the 3.6 35a3 (MOE) model, but for the sorts of tasks I've used Qwen 3.6 to accomplish, the MOE version has always done a great job, and at many times the speed. What that means in practice is that I can perform more iterations in less time, and craft better output overall, using the MOE model. Honestly, I haven't needed any better quality than I get from the MOE model. It's fantastic. When I need more world knowledge and capability, I use Deepseek V4 Flash, and for vision, Mimo 2.5. The Gemma 4 models, Hy3, and Stepfun 3.7 Flash also get thrown into the local AI mix. Orchestrating tasks with that stable of brains has been effective for such a wide variety of workflows.

One thing I think important to note is that Qwen 3.5 122b (MOE) performs more than twice as fast as 3.8 27b (dense) on both Strix Halo and DGX Spark machines. That bigger class of model has much more world knowledge than 27 billion parameters can ever contain. So, I'm very eager to see if Alibaba will release a 120ish billion parameter version of 3.8. I have a sense that a new 100b+ sized MOE model from Qwen could likely be a Deepseek V4 Flash killer.

So for now, I'm disappointed with Qwen 3.8 27b. I'll be extremely eager to try any mixture of experts Qwen 3.8 model. I expect a 35b or 122b size MOE model will likely be a serious improvement over other comparable options. I'll also try other harnesses to see if fewer task termination issues are experienced - and/or perhaps an update to Pi will help reduce the problems I'm currently seeing with 3.8 27b. Either way, I think we'll find that in practice, 3.8 27b is not the magic pill so many people hoped it would be. I'll keep my ears open, and continue testing...

Nick Antonaccio
Nick AntonaccioAdmin
Aug 18, 2026 at 02:30 (edited, 1 revision)
#2

Here's a video about the effectiveness of MTP on 3.8 27b:

https://www.youtube.com/watch?v=NjfHqiNHTxk

Nick Antonaccio
Nick AntonaccioAdmin
Aug 19, 2026 at 12:06 (edited, 3 revisions)
#3

It appears that default settings in LM Studio are not set appropriately for Qwen 3.8 27b. I expect that's the reason for my initial troubles with it. The model card recommends the following sampling parameters:

Thinking Mode (Default):

  • Temperature: 1.0 (LM Studio set it at .1)
  • Top-p: 0.95
  • Top-k: 20 (LM Studio set it at 40)
  • Presence Penalty: 0.0
  • Min_p: 0.0
  • Repetition_penalty: 1.0

Instruct / Non-Thinking:

  • ModeTemperature: 0.7
  • Top-p: 0.80
  • Top-k: 20
  • Presence Penalty: 1.5
  • Min_p: 0.0
  • Repetition_penalty: 1.0
Nick Antonaccio
Nick AntonaccioAdmin
Aug 19, 2026 at 12:08 (edited, 3 revisions)
#4

For GPUs with 12-16Gb VRAM, try the Ridge quant by Empero. It's an architecture-aware quantization averaging 3.7 bits per weight.

It reduces file size to roughly 11.7Gb to 12.6Gb. I'm getting it to run at 20.69 tokens per second without MPT yet, with very good quality results, on the DGX Spark, and 17.47 tokens per second on the Strix Halo laptops.

Things are looking up for 3.8 27b!

Nick Antonaccio
Nick AntonaccioAdmin
Sep 11, 2026 at 15:05 (edited, 2 revisions)
#5

The truth is, so far, I genuinely dislike Qwen 3.8 27b. I know that's an unpopular opinion, but the proof has been in the pudding, in my experiences working with it.

I've run into more 'Response was truncated before completion' messages with it than any other popular model. It's taken so much tuning to get working somewhat reliably, on every machine I own, even the DGX Sparks and the Strix Halos. And I'm always still waiting for when it will fail on a task. The opposite is true of all the MOE models. Qwen 3.6 35a3 just competently zips along on every one of my 9 GPU servers, even the very old machines with 16Gb VRAM (q4 quants run great on those old, still relatively inexpensive to buy used machines).

Qwen 3.8 27b is certainly the smartest model that can be run on 'inexpensive' consumer hardware with low VRAM, but I'd prefer in every case to work with a faster MOE. Qwen 3.6 35a3 has served far better for my needs than 3.8 27b ever has. Perhaps it makes sense to put the time into configuring and tuning 3.8 27b on your hardware, if your goal is get the best possible first-shot visual design & layout results, or if you have absolutely no software development skills and just need a model to do everything for you. But with a little skill, and some time spent iterating through issues (a routine which everyone at the moment seems to be hell-bent on eliminating), I've gotten better, more reliable results from other models which can run on the same hardware. Do we really need to completely remove human input from every single task? The LLMs I use remove so much of the labor and time involved in completing virtually every development task, that software and IT work no longer feel like work at all - even though I still fully expect to do lots of iterations on every project I start. It seems that people want to just say a few words and not have to even look at anything but a first response from an agent. I don't have a desire to work that way. I want to be involved in every process, and use LLMs to take care of the overwhelming majority of tedious, time consuming work. So many models are already passing that bar. whatever single capability you think any 1 particular model offers, is likely not a real requirement. More skill and understanding is the requirement, and if we want human input to continue to be valuable, I think that continuing to develop skill, should be a requirement for anyone using AI tools.

If you have more powerful hardware, go with Deepseek V4 Flash, or even better GLM 5.3 Flash, if you've got the firepower.

It's becoming clear that hardware with 200Gb+ VRAM is what's required to really replace current frontier subscriptions, and there's still quite a gap between self-hosted models and the best API offerings. 512Gb VRAM closes a lot of that gap, but yet, still not entirely, when compared to Fable, Astra, and whatever comes next.

So I still use my $20 ChatGPT routine for all my most demanding software development projects, and I still never pay more than $20 per month. I would reach rate limits very quickly on the API and with Codex, but still never, in the ChatGPT interface. If/when that subsidized use disappears, I'm quite happy with GLM 5.3 Flash IQ3_XXS/IQ4_Xs and/or Deepseek Antirez/DS4 Q4 on two DGX Sparks. That hardware still can be gotten for $10,000ish, which is affordable enough for small-medium size businesses which require totally private self-hosted inference.

Qwen 3.8 Flash Next is their really exciting contender. That architecture seems to be far more effective, but the trained model that came with it is just a preview. I expect the Qwen 4 series will provide a much more solid foundation - let's hope they give us something that runs on a single machine with 128Gb VRAM.

I'm also very eager to see what Poolside delivers after the $7B commitment they got from Nvidia. Laguna S2.1 already runs at over 30 tokens per second on a single Strix Halo machine, and it's currently got a 1 million token context limit (compared to Qwen 3.8 27b's 256,000 limit). With real money and support from Nvidia, I expect Poolside will give us another new baseline for single 128Gb machines which people and small businesses can afford to buy without too much financial pain. Keep your eye on their next releases.

The last quarter of 2026 should be very exciting for self-hosting!

Please login to post a reply.

© 2026 AI By Nick.