Qwen3.8-27B on Strix Halo — Final Results
Hardware: ASUS ROG Flow Z13 GZ302EA · AMD Ryzen AI Max+ 395 · Radeon 8060S (gfx1151) ·
128 GB unified LPDDR5X with 96 GB carved to the iGPU (Windows sees 31.6 GB)
Software: LM Studio 0.4.24+1 · llama.cpp ROCm 2.34.0 engine · Windows 11 (build 26200)
Test: 200 generated tokens, 8k context, single stream (--parallel 1), temp 0.7, clean boot
TL;DR — the two configs to use
| Goal | Model | Load flags | Speed |
|---|---|---|---:|
| Best 4-bit for speed | qwen/qwen3.8-27b (lmstudio-community Q4_K_M) | --gpu max --parallel 1 --speculative-draft-mtp --speculative-draft-max-tokens 2 | ~18 tok/s |
| Best 8-bit for quality | qwen3.8-27b@q8_k_xl (unsloth UD-Q8_K_XL) | same flags as above | ~10 tok/s |
Both are lossless: speculative decoding verifies every drafted token with the full model, so
output quality is identical to running without it. MTP (multi-token prediction) heads are bundled in these GGUF files and are enabled by the --speculative-draft-mtp flag.
Copy-paste (LM Studio CLI)
LMS=~/.lmstudio/bin/lms.exe
# Best 4-bit for speed (~18 tok/s)
"$LMS" load "qwen/qwen3.8-27b" --gpu max --parallel 1 --context-length 8192 \
--speculative-draft-mtp --speculative-draft-max-tokens 2
# Best 8-bit for quality (~10 tok/s)
"$LMS" load "qwen3.8-27b@q8_k_xl" --gpu max --parallel 1 --context-length 8192 \
--speculative-draft-mtp --speculative-draft-max-tokens 2
Measured results (clean boot, ROCm engine, parallel 1)
| Configuration | tok/s | Draft acceptance | |---|---:|---:| | Q4_K_M, MTP off | 11.7 | — | | Q4_K_M + MTP, draft depth 2 | 18.1 | 108/181 (60%) | | Q4_K_M + MTP, LM Studio auto depth | 17.5 | 120/235 | | Q4_K_M + MTP, draft depth 3 | 15.7 | 107/273 | | UD-Q4_K_XL + MTP, draft depth 3 | 16.5 | 113/256 | | UD-Q8_K_XL + MTP, draft depth 2 | 10.2 | 103/191 (54%) | | Q4_K_M + MTP off (Vulkan engine, for reference) | 14.6 | — |
Takeaways
- MTP is the single biggest free win: +55% on Q4_K_M (11.7 → 18.1 tok/s). It requires no download or fork for these files — the heads ship inside the GGUFs.
- Draft depth 2 is optimal on this bandwidth-bound iGPU. Depth 3+ wastes bandwidth on rejected drafts and collapses throughput (depth 4 was catastrophic in testing).
- Use the ROCm engine, not Vulkan (+18% on the same model/flags).
- The official "Recommended" quant (Q4_K_M) is also the fastest — no need to go higher for speed. Going to 8-bit costs ~1.8× throughput for near-fp16 fidelity.
- The unsloth
UD-Q4_K_XLoften recommended for other GPUs is slower here (16.5 tok/s); its dynamic mixed-quant tensors don't suit the ROCm k-quant kernels. Skip it.
One critical caveat - loading which requires reboot
Windows only sees 31.6 GB of the 128 GB because 96 GB is reserved for the iGPU. Every large model load fills the Windows standby cache. After ~5–6 loads it reached 19.2 GB and decode collapsed ~4× across all models — same config that ran 18.1 tok/s dropped to 4.2 tok/s — with no error and no health-check warning. A reboot clears it and restores full speed.
Practical rules:
- Benchmarking: reboot first, run one paired A/B test, then stop. Don't chain many model loads.
- Daily use: keep the model you're using loaded instead of cycling several large models.
- If throughput mysteriously tanks: reboot. It is memory-cache pressure, not the model.
Also: --parallel 1 matters. Multi-slot servers depress the baseline (~20%) and speculative
decoding only helps single-stream requests.
What did NOT work
- kingjones777 "ROCmFP4-STRIX-MTP" quant (separate MTP head): does not load in stock LM Studio — it uses ggml tensor types 100–106 that only exist in the
ROCmFPXllama.cpp fork. Chasing it means installing the ROCm/HIP SDK, building the fork from source, and running a standalonellama-server. Potential reward 24–45 tok/s, but it leaves LM Studio entirely. - External small draft model (Qwen3.5-0.8B as a "simple" draft): only ~25% acceptance and slower (4.0 tok/s). Use the bundled MTP head instead.
Recommended settings summary
| Setting | Value |
|---|---|
| Engine | llama.cpp-win-x86_64-amd-rocm-avx2 (latest, 2.34.0) |
| GPU offload | max (100%) |
| Parallel / concurrency | 1 |
| Context length | 8192 (raise freely; 96 GB pool allows 64k+) |
| Speculative decoding | MTP on, max-tokens 2 |
| Sampler (for latency) | Lower temperature than the 1.0/top_p 0.95 default; defaults make the model overthink |
Report generated from live measurements on the machine described above. Numbers are hardware-specific; expect different values on different GPUs.