Deepseek V4 then V4.1 Flash have been my workhorse models for months, both locally & on the APIs (but recently, GLM 5.3f, Mimo 2.6f, & Qwen 3.8f Next have been great alternatives for local hosting)

1296 views Pinned
Nick Antonaccio
Nick AntonaccioAdmin
Aug 01, 2026 at 02:27
#11

DS4 Flash is now out preview:

https://www.youtube.com/watch?v=PTdu0JlhGfw

Nick Antonaccio
Nick AntonaccioAdmin
Aug 02, 2026 at 15:11 (edited, 4 revisions)
#12

The new version of V4 Flash out of preview (0731) is going viral, because it scores even higher than V4 Pro on many benchmarks, and users are seeing it produce fantastic quality output. It's still the same model architecture, but they apparently improved the post-training data quality significantly (the vast set of devised question-answer pairs which train the model how to act, and sculpt the sorts of results it delivers).

The 0731 build is live on deepseek.com and Openrouter now, and the open weights have been released. The open source versions are a drop-in replacement for the preview versions, in the ds4 framework, for those who are self hosting.

The 2-8 mixed bit quant still runs well and does a very good job, for all the sorts of tasks I rely on, when served on a single machine with 128Gb VRAM. The 4 bit quant does even better job, with more room for full length KV cache, on a cluster of 2 of the same machines (DGX Sparks, Strix Halos, Macs with 128Gb, 2 RTX 6000s, etc.).

I expect Deepseek V4 Flash on 2 Asus GX10 machines to be the default foundation for self-hosted inference, which I'll recommend to clients. Qwen 3.6, Gemma 4, Mimo 2.5, and Hy3, are all still great alternatives for fast sub-task work, but V4 Flash will be my local LLM backbone. It's proven itself to be rock solid in all my work.

V4 Flash has honestly come to dominate a large portion of all my production LLM use. I've been using it to complete virtually every IT-related task, since the day I began using it, and I've never seen it fumble. It's ridiculously fast, extraordinarily capable for its size, and outrageously cheap on the Cline-pass provider, plus there are plenty of backup providers, including Openrouter, deepseek.com, Ollama, and elsewhere. I've used V4 Flash to build an absolutely huge variety of small-medium sized utility applications, and many significant portions of code bases that fit into larger production software stacks. It's been a rock solid reliable developer, in the Pi coding harness (I should make it clear that my ChatGPT zip file development routine still forms the foundation of most large-scale production work for my big clients, and it will continue to be my go-to LLM for large Python/Flask/Web apps, for as long as OpenAI provides apparently endless rate limits for frontier use, at $20 per month - but I am no longer frightened of the moment they go out of business).

V4 Flash has absolutely become my go-to model for any significant local inference work. I perform lots of tasks which involve PHI (health data that requires strict HIPAA compliance to avoid legal repercussions), which can't be entered into typical public LLM API providers such as Openrouter - or which is very expensive to process on HIPAA compliant providers such as Bastion. V4 Flash is my go-to for any challenging tasks in that domain.

I still use Qwen 3.6 often for simple local tasks that must be compliant, because it performs so much faster than V4 Flash, but if there's any question about the potential for a less than acceptable quality of work by Qwen, V4 Flash gets the job from the get-go. I wrote a bit about one such experience in this topic: https://aibynick.com/thread/60

I've gotten V4 Flash performing at 13-27 tokens per second locally, depending on which version I'm running and other environment factors. That's plenty fast for most of the batch jobs I use it to complete, so I tend to skip even trying smaller models most of the time.

I did see one reviewer claim that V4 Flash hallucinates certain classes of work more than other similarly sized models: https://www.youtube.com/watch?v=iT34JqYtGRs , so he prefers to rely on Hy3 (but that same person also made this glowing review: https://www.youtube.com/watch?v=wT42SgaOPK4 , so take his contrary results with a grain of salt).

The majority of other initial V4 Flash reviews have been dramatically good. There are so many fantastic examples of all kinds, where V4 Flash's output beats the quality of all but the very best most recent frontier models (it even beats Kimi K3 in low reasoning mode, in some benchmark results (what?!?!? holy moly that's impressive)), and its amazing benchmark improvements compared to the preview release support the likelihood that version 0731 competes with frontier models in many domains of work.

V4 Flash provides a perfect mix of capability, performance, and self-hostable size, to make it a smart fit for an enormous scope of work. It enables you to run the same model on an API for speed (virtually for free on a provider like Cline-pass), and locally for privacy and self-reliance. The consistency that comes from using the exact same model for so many tasks, in so many situations, makes a big difference in how you work. You do become accustomed to expected workflow patterns, verbiage, strengths/weaknesses, etc., with a model that you know intimately and have used repeatedly. The experiences you work through, while building solutions with the same model, form a baseline workflow which is similar to personal familiarity you gain while working with colleagues over long term development projects. That familiarity breeds trust, reliability, and improved productivity. Switching workflows and expectations, testing results, etc., takes up an enormous amount of time, when dipping your feet into production work with a new model.

I expect V4 Flash will shape the industry in practical directions which go well beyond just constantly competing to build the next bigger multi-trillion parameter frontier model. New levels of frontier intelligence will certainly continue to shape the future, but most of us just need an affordable, fast, and capable model like V4 Flash to get the overwhelming majority of daily work completed. We need more of this sort of model: smaller, faster, cheaper, smarter for its size, and genuinely usable on modest self-hosted hardware.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 08, 2026 at 22:20 (edited, 1 revision)
#13

I thought it might be useful to compare some output from the full uncompressed version of ds4f (on cline-pass in this case):

https://com-pute.com/nick/rubiks_cube.html

and exported Pi session in which that app was built:

https://com-pute.com/nick/rubiks_cube_cline_ds4f--pi-session-2026-08-08T13-47-55-212Z_019fe1a1-57cc-7882-8259-9b7cca666a93.html

To a version of the same app, created with the 2 bit compressed version that runs in the ds4 engine, on a single DGX Spark (Asus GX10):

https://com-pute.com/nick/rubiks--dsf4-local.html

and the Pi session in which it was created:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

I love that this tiny box which sits on the floor (and could fit in a handbag), can reliably write working code like this. And I love that I'm able to use the same model locally that I use on the cline-pass API. I'm also fully aware of how good a buy that cline-pass API is. I use the API version all day every day, on an account that costs less than $7 per month.

You can see the difference in quality between the quantized local version and the uncompressed version running on cline-pass. If you take a brief look at the session which used the locally hosted quantized version, the model required many more iterations, needed guidance completing the task, and did not create as nice of a final application as the uncompressed version on the API (more features were added to the app created by the uncompressed LLM, and the UI looked better in that app).

Quantized versions of models are like drunk versions of themselves. They have the same background as their full precision versions, but they make more mistakes in judgement and have trouble thinking things through as deeply.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 09, 2026 at 03:53
#14

BTW, Laguna S2.1 was total junk on this same prompt:

https://aibynick.com/thread/57?page=1#post-181

Surprisingly, Qwen 3.6 35a3 did a great job, in just a few minutes:

https://com-pute.com/nick/qwenrubiks.html

Nick Antonaccio
Nick AntonaccioAdmin
Aug 09, 2026 at 18:55
#15

I performed a comparison between Hy3, Deepseek V4 Flash, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver. I think Hy3 did the best job out of the gate:

Hy3 got it done quickly, with the fewest iterations and issues, and it built a nice UI. Here's the Pi session export:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

I've got to note that Mimo 2.5 also did a great job, very quickly:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

Nick Antonaccio
Nick AntonaccioAdmin
Aug 19, 2026 at 12:10 (edited, 2 revisions)
#16

DeepSeek is increasing its API token prices between 50%-1,100%.

They're switching to a dynamic peak and off-peak billing model. Peak hours are designated as 01:00–04:00 UTC and 06:00–10:00 UTC, during which rates are exactly double the new off-peak prices. Pricing updates take effect at 16:00 UTC on August 16, 2026.

Here's a breakdown per 1 million tokens:

DeepSeek Model & Token Type Old Price New Off-Peak Price New Peak Price Max%

V4-Flash Output $0.2800 $0.6600 $1.3200 +371%

V4-Flash Input (Cache Miss) $0.1400 $0.2200 $0.4400 +214%

V4-Flash Input (Cache Hit) $0.0028 $0.0070 $0.0140 +400%

V4-Pro Output $0.8700 $1.9800 $3.9600 +355%

V4-Pro Input (Cache Miss) $0.4350 $0.6600 $1.3200 +203%

V4-Pro Input (Cache Hit) $0.0036 $0.0220 $0.0440 +1114%

Nick Antonaccio
Nick AntonaccioAdmin
Aug 16, 2026 at 15:41 (edited, 2 revisions)
#17

Be aware that Deepseek pricing is now doubled during peak hours, so for the East Coast USA where I am, that means avoid nighttime use. Specifically, prices are doubled during these periods:

  • 9:00pm - midnight
  • 2:00am - 6:00am

Said the other way, use Deepseek during the day:

  • 6am - 9pm (and a little window midnight - 2am)
Nick Antonaccio
Nick AntonaccioAdmin
Aug 21, 2026 at 10:19
#18

Something is wrong with Deepseek Flash 0731. I just tried a few simple chat questions with it in Jan, using OpenRouter. The old 0423 version worked without any issues. 0731 argued with itself during a much longer thinking process, and then produced less incorrect output.

Nick Antonaccio
Nick AntonaccioAdmin
Aug 28, 2026 at 14:08
#19

Lemonade 11.8 Makes It Easy To Run DeepSeek V4 Flash On AMD Strix Halo:

https://www.phoronix.com/news/AMD-Lemonade-11.8

Nick Antonaccio
Nick AntonaccioAdmin
Aug 30, 2026 at 11:53
#20

There have been some issues with the thinking blocks in Pi printing out 1 word per line, when using V4 Flash and V4 Pro on the Openrouter and Cline-pass providers. This appears to be an issue with the providers' output, rather than with Pi, and it doesn't appear to stop the model from thinking properly - just makes for vertically lengthy console printouts. If it bothers you, switching to the Deepseek API eliminates the issue.

Please login to post a reply.

© 2026 AI By Nick.