Laguna S 2.1 is the newest great mid-size self-hostable model.

64 views
Nick Antonaccio
Nick AntonaccioAdmin
Jul 23, 2026 at 23:44 (edited, 1 revision)
#1

Laguna S 2.1 has been made to be specifically capable as a coding model. It's another entry in the class of models such as Deepseek V4 Flash and Hy3, which runs on a single DGX Spark or Strix Halo machine, with plenty of VRAM to spare for KV Cache. At Q4_K_M quant it's far smaller than either Hy3 at Q3 or V4 Flash at Q2 , and it comes in plain old GGUF format, so it doesn't need any sort of special framework to run (such as DS4 for Deepseek). Initial tests with it show really fantastic coding performance - I'll be putting it through the paces over the next week. This Youtuber provides some initial examples:

https://www.youtube.com/watch?v=ERCdFgTXNY4

It's also great to a US company providing some competitive self-hostable models.

Nick Antonaccio
Nick AntonaccioAdmin
Jul 27, 2026 at 15:20
#2

I see a bit of a consensus forming around Laguna S 2.1: it appears to be a perfectionist. It's a capable model, but often gets into a loop evaluating it's own decisions, and never getting the work done, where less capable models produce some less than perfect output which you can choose to iterate upon.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 07, 2026 at 14:27
#3

I've had trouble running Laguna in LM Studio on DGX Spark, but on the Strix Halo machines it runs at 32 tokens per second at q4km quant. So far, no other model has dethroned Deepseek v4 Flash for local inference on 128Gb class machines.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 09, 2026 at 03:55 (edited, 2 revisions)
#4

Laguna S2.1 has been a significant disappointment. Here's a version of a Rubik's Cube solver that Laguna S2.1 (q4km quant) worked on for several hours, and never got working completely:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail.html

Here's the Pi session:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Deepseek v4 Flash did a great job with the same prompt:

https://aibynick.com/thread/50?page=2#post-180

Even Qwen 3.6 35a3 did a great job - and in just a few minutes on a local machine (and this version of Qwen can run on a much smaller GPU):

https://com-pute.com/nick/qwenrubiks.html

Please login to post a reply.

© 2026 AI By Nick.