End of September 2026 recommendations are here: https://aibynick.com/thread/80
Many valid details are also still available in the August-mid September 2026 roundup: https://aibynick.com/thread/63
The way it currently stands, for the overwhelming majority of daily work, I use Deepseek V4.1 Flash on the Cline-pass API. That account, which costs less than $7 per month, is one of the most valuable services in my life. I build most software and complete IT work (installations, OS configs, compilations, etc.) with it all day every day.
For any task that involves PHI or compliance obligations, that's where my self-hosted models are required. I run Mimo 2.6 Flash MXFP4 on a cluster of 2 DGX Sparks, and GLM 5.3 flash iq3_xxs on a separate cluster of 2 DGX sparks. Those jobs most often involve working with private account credentials, building software which can't be shared publicly, and anonymizing data in various file formats (so I can work with the shape of that data in faster APIs such as Cline-pass).
Qwen 3.8 Flash Next iq4_xs currently tends to stay loaded on my Strix Halo machines. It's amazingly capable, fast, full of world knowledge, able to understand and reason through challenging problems quickly, etc. To be honest, I could likely do everything I need with only that model, if I had to. GLM and Mimi are even more capable, so they require less hand holding and iterations, but with a little elbow grease, QFN is capable of replacing lots of infrastructure and expenses.
Qwen 3.6 35a3 q4 runs most often on my 3 machines which have RTX GPUs with 16-24GB VRAM (that old model is fantastically capable, for a tiny, incredibly fast model).
I also keep a Deepseek API key handy for quick configs and for use in temporary environments where Cline-pass takes longer to set up, and I use OpenRouter to test new models.
For the overwhelming majority of my daily work, I run every model in Pi coding agent, on local machines (even on my Android phone, in Termux), and on several VPS account hosted at Contabo ($5 per month). Locally, I run it on the command line. On the VPSs, I use the Pi-web interface, which lets me connect and run it in any browser, on any device with Internet access. Because of Pi-web, I've begun to use the Pi installation in the VPS account most often.
I can use any of the models, from any of the LLM API provider accounts, or any of my self hosted models, in any of the Pi installations - though Deepseek v4.1 Flash is really all I default to constantly these days.
It's finally getting to the point that that combination of ds41f model, pi harness, pi-web UI, LLM API, and VPS hosting are replacing all my other workflows, including the ChatGPT zip file workflow which has been a cornerstone of my development work for more than a year (I still keep all the legacy projects in ChatGPT, but the new pi-web environment works better for new projects, and it's no biggie to switch and ChatGPT project over to pi-web, if needed).
To manage a fleet of old 32-bit Windows 7 machines, I compiled a version of Nullclaw (https://aibynick.com/thread/88). That harness makes those ancient machines capable of completing all sorts of useful tasks, with the same effortless agentic workflow we've gotten used to using on new machines. Since RAM and other components are currently so expensive, those old machines can be a solution for people and organizations who don't have piles of money to spend on hardware. $80 per year Cline-pass account is all you need.