Post History

Current version by Nick Antonaccio

Current VersionOct 08, 2026 at 23:58

I've been keeping a copy of Mimo 2.6 Flash running on a cluster of 2 DGX Sparks. It's quicker than GLM 5.3 Flash, and around the same class of capability. I keep GLM 5.3 Flash running on a separate cluster of 2 DGX Sparks, and use them both interchangeably.

Previous Versions
Version 7Oct 08, 2026 at 23:58

I've been keeping a copy of Mimo 2.6 Flash running on a cluster of 2 DGX Sparks. It's quicker than GLM 5.3 Flash, and around the same class of capability. I use them both.

Version 6Oct 08, 2026 at 23:57

I've been keeping a copy of Mimo 2.6 Flash running on a cluster of 2 DGX Sparks. It's quicker than GLM 5.3 Flash, and around the same class of capability.

Version 5Oct 08, 2026 at 23:56

The way it currently stands, for the overwhelming majority of daily work, I actually use Deepseek V4.1 Flash on the Cline-pass API. That $7 per month account is one of the most valuable services in my life. I build most software, complete IT work (installations, OS configs, compilations, etc.), with it all day every day.

For any task that involve PHI and compliance obligations, I run Mimo 2.6 Flash MXFP4 on a cluster of 2 DGX Sparks, or GLM 5.3 flash iq3_xxs on a separate cluster of 2 DGX sparks. Those jobs most often involve anonymizing data in various file formats, working with private account credentials, and building software which can't be shared publicly.

Qwen 3.8 Flash Next iq4_xs tends to stay loaded on my Strix Halo machines. It's amazingly capable, fast, full of world knowledge, able to understand reason through problems quickly, etc. I could honestly likely do everything I need with only that model, if I had to. GLM and Mimi are just even more capable, and require less hand holding.

Qwen 3.6 35a3 q4 runs most often on my 3 machines which have RTX GPUs with 16-24GB VRAM (it's fantastic capable, for a tiny, incredibly fast model).

I also keep a Deepseek API key handy for quick configs, and use OpenRouter to test new models.

I run every model in Pi coding agent, on local machines, and on a VPS at Contabo ($5 per month), via the Pi-web interface (in any browser, on any device connected to the Internet). I use the Pi installation in the VPS account most often.

I can use any of the models on any of the API accounts, or any of my self hosted models, in any of the Pi installations, though Deepseek v4.1 Flash is really all I default to constantly these days.

Version 4Oct 08, 2026 at 23:47

The way it currently stands, for the overwhelming majority of daily work, I actually use Deepseek V4.1 Flash on the Cline-pass API. That $7 per month account is one of the best things in my life. I build most software, complete IT work (installations, OS configs, compilations, etc.), with it all day every day.

For any task that involve PHI and compliance obligations, I run Mimo 2.6 Flash MXFP4 on a cluster of 2 DGX Sparks, or GLM 5.3 flash iq3_xxs on a separate cluster of 2 DGX sparks. Those jobs most often involve anonymizing data in various file formats, working with private account credentials, and building software which can't be shared publicly.

Qwen 3.8 Flash Next iq4_xs tends to stay loaded on my Strix Halo machines. It's amazingly capable, fast, full of world knowledge, able to understand reason through problems quickly, etc. I could honestly likely do everything I need with only that model, if I had to. GLM and Mimi are just even more capable, and require less hand holding.

Qwen 3.6 35a3 q4 runs most often on my 3 machines which have RTX GPUs with 16-24GB VRAM (it's fantastic capable, for a tiny, incredibly fast model).

I also keep a Deepseek API key handy for quick configs, and use OpenRouter to test new models.

I run every model in Pi coding agent, on local machines, and on a VPS at Contabo ($5 per month), via the Pi-web interface (in any browser, on any device connected to the Internet). I use the Pi installation in the VPS account most often.

I can use any of the models on any of the API accounts, or any of my self hosted models, in any of the Pi installations, though Deepseek v4.1 Flash is really all I default to constantly these days.

Version 3Oct 08, 2026 at 22:27

The way it currently stands, I run Mimo 2.6 Flash MXFP4 on a cluster of 2 DGX Sparks, GLM 5.3 flash iq3_xxs on a separate cluster of 2 DGX sparks. They're both great.

Qwen 3.8 Flash Next iq4_xs tends to stay loaded on my Strix Halo machines. It's amazingly capable (I could probably do everything I need with only that model).

Qwen 3.6 35a3 q4 runs most often on the 3 machines with RTX GPUs that have 16-24GB VRAM (it's fantastic for a tiny, incredibly fast model).

For the overwhelming majority of daily work, I actually use Deepseek V4.1 Flash on the Cline-pass API. That $7 per month account is one of the best things in my like.

I also keep a Deepseek API key handy for quick configs, and use OpenRouter to test new models.

I run every model in Pi coding agent, on local machines, and on a VPS at Contabo ($5 per month), via the Pi-web interface (in any browser, on any device connected to the Internet). I use the Pi installation in the VPS account most often.

Version 2Oct 08, 2026 at 22:19

The way it currently stands, I run GLM 5.3 flash iq3_xxs on a cluster of 2 DGX Sparks, Mimo 2.6 Flash MXFP4 on a separate cluster of 2 DGX sparks. Qwen 3.8 Flash Next iq4_xs tends to stay loaded on my Strix Halo machines, and Qwen 3.6 35a3 q4 runs most often on the 3 machines with RTX GPUs that have 16-24GB VRAM.

For the overwhelming majority of daily work, I use Deepseek V4.1 Flash on the Cline-pass API ($7 per month). I keep a Deepseek API key handy for quick configs, and use Openrouter to test new models.

I run every model in Pi coding agent, on local machines, and on a VPS at Contabo ($5 per month), via the Pi-web interface (in my browser, on any device). I use the Pi installation in the VPS account most often.

Version 1Oct 08, 2026 at 22:15

The way it currently stands, I run GLM 5.3 flash on a cluster of 2 DGX Sparks, Mimo 2.6 Flash on a separate cluster of 2 DGX sparks. Qwen 3.8 Flash Next tends to stay loaded on my Strix Halo machines, and Qwen 3.6 35a3 runs most often on the 3 machines with GPUs that have 16-24GB VRAM.