The way it currently stands, for the overwhelming majority of daily work, I actually use Deepseek V4.1 Flash on the Cline-pass API. That $7 per month account is one of the most valuable services in my life. I build most software, complete IT work (installations, OS configs, compilations, etc.), with it all day every day.
For any task that involve PHI and compliance obligations, I run Mimo 2.6 Flash MXFP4 on a cluster of 2 DGX Sparks, or GLM 5.3 flash iq3_xxs on a separate cluster of 2 DGX sparks. Those jobs most often involve anonymizing data in various file formats, working with private account credentials, and building software which can't be shared publicly.
Qwen 3.8 Flash Next iq4_xs tends to stay loaded on my Strix Halo machines. It's amazingly capable, fast, full of world knowledge, able to understand reason through problems quickly, etc. I could honestly likely do everything I need with only that model, if I had to. GLM and Mimi are just even more capable, and require less hand holding.
Qwen 3.6 35a3 q4 runs most often on my 3 machines which have RTX GPUs with 16-24GB VRAM (it's fantastic capable, for a tiny, incredibly fast model).
I also keep a Deepseek API key handy for quick configs, and use OpenRouter to test new models.
I run every model in Pi coding agent, on local machines, and on a VPS at Contabo ($5 per month), via the Pi-web interface (in any browser, on any device connected to the Internet). I use the Pi installation in the VPS account most often.
I can use any of the models on any of the API accounts, or any of my self hosted models, in any of the Pi installations, though Deepseek v4.1 Flash is really all I default to constantly these days.