Post History

Current version by Nick Antonaccio

Current VersionSep 28, 2026 at 13:46

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI development usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI ever imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

Mimo 2.6 Pro and Flash have also been added to Cline-pass, and they're quickly becoming another set of preferred models. MiMo-V2.6-Pro achieved a score of 46 on the Artificial Analysis Intelligence Index. That's the highest score for any open-weight model yet - equivalent to Opus 5 and GPT-5.6 Sol. At the same time, it's also rated as the cheapest model at $0.13 per task. And don't forget the Flash model at $0.14/$0.28 in/out. These are serious contenders. My initial tests of both Mimo 2.6 models on Openrouter have proved them to be as impressive in practice as they are in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you leverage LLM capabilities to engineer your workflows. Your LLMs should regularly be used to build the tooling which your LLMs use in project. Pi is built around simplifying that approach to solving problems. That's one of the main reasons Pi has become my most used harness - it's made to enable editing & extending of not only skills and tools, but also the code of the harness itself. But don't fall into the trap of thinking any particular harness is required to get work completed - you can use any harness to write code to build tools and system components. More important than any set of built-in features, harnesses simply give an LLM hands to work with files and local operating system operations. It's up to you to envision, specify, and direct the LLM's capabilities to build whatever tooling is required to solve a problem. Relying on the tooling and skills built into heavy harnesses can certainly be a time saver if you don't know how to build those things yourself, but that reliance also tends to hide the nature of what makes LLM driven development so fantastic - that the LLMs can introspect and build out their own tooling, if you just know to ask.

Codex is the next most interesting harness of the current crop, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Pi-web

I absolutely love using pi-web as an interface to Pi in my browsers:

https://aibynick.com/thread/86

Pi-web is becoming a permanent, prominent addition to all my AI tools, and may very well become the primary interface I use for all my workflows with LLMs. It gives me access to all my models (self-hosted and on the APIs), using my own application servers, with all of my existing files & sessions, with all the same controls Pi gives me on the command line, but in any browser, on any device, from anywhere the Internet is available - and it takes about a minute to set up.

Pi-web is not a new piece of software this month, but it's new to me, so I've included it this update. It's one of the most worth-while recent tools to check out.

Locally Hosted Models

My favorite locally hosted models are currently:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context), Mimo 2.6 Flash (at native MXFP4 for its MoE experts, with full 1M context), and Deepseek v4 Flash (2-4 bit mixed quant, 1M context, running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated (think, what an office staff would use to perform tasks). You can always add more DGX Spark machines later to handle bigger models, more users, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next (IQ4_XS and IQ3_XXS quants, depending on required context length), the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Gemma 4 models are also useful on small GPUs. Perhaps consider Qwen 3.8 27b as an orchestrator model in your mix, but have the smaller MOEs do as much work as possible. In general, on smaller GPUs, you'll need to employ several models, and use the best one for each task, with a larger model checking the work of smaller models, and fixing issues which the smaller models fail on.

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is an exceptionally capable coder, has a lot of general world knowledge, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) which will be relatively affordable for at least the next few years.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Be sure to see how much better Qwen 3.8 Flash Next did in the comparison of models building Rubik's Cube demos, at https://aibynick.com/thread/54#post-185 .

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less (https://www.amazon.com/gp/product/B0GWK5MPJ5/ref=ox_sc_saved_image_1). The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is currently hard to beat, for the price and power requirements.

The 256Gb Apple M5 Ultra machines were just released 9-22-2026, and in initial reviews, they appear to be far more performant than the M3 Ultras. I haven't had a chance to use them yet. Those boxes are the most energy efficient computers available for self-hosting very large models - and after the 512Gb RAM models are released in late October 2026, I may be tempted to get a 2 machine cluster. If pre-fill and other performance reviews shake out as expected, there won't be much simpler hardware available for running huge terabyte parameter models locally.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. Interestingly, my initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful on Strix Halo (that lobotomized low quant ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs or an old laptop with a mobile RTX 3080ti, and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can absolutely end up be being a hassle, with lots of hurdles to jump over along the way.

The lowest price solutions I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is quite questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately effective overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor.

Learning to work with ternary models can actually yield some productive capability, even on sub-$100 netbooks (that's what I was using in the linked post above), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the best test outcomes for me. Just don't get caught up in any expectations that ternary models will remotely replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Previous Versions
Version 26Sep 28, 2026 at 13:46

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI development usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI ever imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

Mimo 2.6 Pro and Flash have also been added to Cline-pass, and they're quickly becoming another set of preferred models. MiMo-V2.6-Pro achieved a score of 46 on the Artificial Analysis Intelligence Index. That's the highest score for any open-weight model yet - equivalent to Opus 5 and GPT-5.6 Sol. At the same time, it's also rated as the cheapest model at $0.13 per task. And don't forget the Flash model at $0.14/$0.28 in/out. These are serious contenders. My initial tests of both Mimo 2.6 models on Openrouter have proved them to be as impressive in practice as they are in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Pi-web

I absolutely love using pi-web as an interface to Pi in my browsers:

https://aibynick.com/thread/86

Pi-web is becoming a permanent, prominent addition to all my AI tools, and may very well become the primary interface I use for all my workflows with LLMs. It gives me access to all my models (self-hosted and on the APIs), using my application servers, with all of my existing files & sessions, and all the same controls Pi gives me on the command line, but in any browser, on any device, from anywhere the Internet is available - and it takes about a minute to set up.

Pi-web is not a new piece of software this month, but it's new to me, so I've included it this update. It's one of the most worth-while recent tools to check out.

Locally Hosted Models

My favorite locally hosted models are currently:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context), Mimo 2.6 Flash (at native MXFP4 formatting for its MoE experts, with full 1M context), and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next (IQ4_XS and IQ3_XXS quants), the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model in your mix, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is an exceptionally capable coder, has a lot of general world knowledge, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Be sure to see how much better Qwen 3.8 Flash Next did in the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185 .

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

The 256Gb Apple M5 Ultra machines were just released 9-22-2026, and in initial reviews, they appear to be far more performant than the M3 Ultras. Those boxes are the most energy efficient computers available for self-hosting very large models - and after the 512Gb RAM models are released in late October 2026, I may be tempted to get a 2 machine cluster. If pre-fill and other performance reviews shake out as expected, there won't be much simpler hardware available for running huge terabyte parameter models locally.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is currently hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (that lobotomized low quant ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs or an old laptop with a mobile RTX 3080ti, and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a hard hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is quite questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post above), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive test workflows for me. Just don't get caught up in any expectations that ternary models can remotely replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Version 25Sep 28, 2026 at 02:10

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI development usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

Mimo 2.6 Pro and Flash also been added to Cline-pass, and they're quickly becoming another set of preferred choices. MiMo-V2.6-Pro achieved a score of 46 on the Artificial Analysis Intelligence Index. That's the highest score for any open-weight model yet - equivalent to Opus 5 and GPT-5.6 Sol. At the same time, it's also rated as the cheapest model at $0.13 per task. And don't forget the Flash model at $0.14/$0.28 in/out. These are serious contenders. My initial tests of both Mimo 2.6 models on Openrouter have proved them to be as impressive in practice as they are in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Locally Hosted Models

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Version 24Sep 28, 2026 at 01:35

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices. MiMo-V2.6-Pro achieved a score of 46 on the Artificial Analysis Intelligence Index. That's the highest score for any open-weight model yet - equivalent to Opus 5 and GPT-5.6 Sol. At the same time, it's also rated as the cheapest model at $0.13 per task. And don't forget the Flash model at $0.14/$0.28 in/out. These are serious contenders. My initial tests of both Mimo 2.6 models on Openrouter have proved them to be as impressive in practice as they are in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Locally Hosted Models

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Version 23Sep 26, 2026 at 15:10

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices, because they're currently the best cost-per-task model rated by Artificial Analysis. My initial tests of Mimo 2.6 on Openrouter have proved it to be as impressive in practice as it is in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Here's my list of harnesses which are currently popular in common use:

Locally Hosted Models

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Version 22Sep 26, 2026 at 15:04

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices, because they're currently the best cost-per-task model rated by Artificial Analysis. My initial tests of Mimo 2.6 on Openrouter have proved it to be as impressive in practice as it is in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Locally Hosted Models

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the State of Affairs article from August-September. The approaches to working with agents have standardized quite a bit: https://aibynick.com/thread/63

Version 21Sep 26, 2026 at 14:56

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

Hosted models

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total expense per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of free world class AI usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

If/when OpenAI imposes rate limits, I'd immediately replace it with additional Cline-pass accounts. 3 monthly Cline-pass accounts can be purchased for about the same price as a single ChatGPT $20 subscription - that's how good a buy Cline-pass is - no other unsubsidized option I'm aware of even comes close to the value of Cline-pass.

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and it's the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is extremely rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been a leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices, because they're currently the best cost-per-task model rated by Artificial Analysis. My initial tests of Mimo 2.6 on Openrouter have proved it to be as impressive in practice as it is in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

Harnesses

I use Pi as my harness nearly universally for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, an auto-loading file such as agents.md, and some optimization routines - by any harness - you can get that harness to complete virtually any required work on your local computer. That provides the basis to prompt your model to build a skill system, install and run libraries, write code to create tools of any sort, control the PC, etc. From what I see, the majority of the public AI market consists of users who expect every agentic capability to work first-shot, out of the box - so when any configuration needs to be tweaked even the tiniest bit, the overwhelming majority of users just don't know what to do, and immediately assume that tasks simply can't be implemented with a given harness. In my experience, that's never the case. You just have to understand how the underlying systems work, and get the model to build the proper tools and configurations.

What matters most is still the raw capability of the LLM(s) you use, your understanding of the systems you use, and the way you engineer your workflows.

Codex harness is most interesting otherwise, for it's built-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and third party computer use extensions work great with Pi.

Locally Hosted Models

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

Local Hardware

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 20Sep 26, 2026 at 14:55

Much of my August-September 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are DeepSeek v4 to v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice, if you agree to share data with Meta), and Qwen 3.8 Flash Next (currently a very promising option for local use on less than $10,000 of hardware).

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish total per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier. If you're not taking advantage of that workflow, you're leaving an absolutely massive amount of world class AI usage untapped.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash and other Flash models (other bigger models will of course use up your account much more quickly).

GLM 5.3 Flash has recently been added to Cline-pass, so it's my preferred choice for vision work, and is the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (such a situation is very rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. Earlier this year, Muse 1.3 would have been the leading frontier model, so don't look past it just because the frontier is moving so quickly. This class of model is a tremendous work horse.

When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices, because they're currently the best cost-per-task model rated by Artificial Analysis. My tests of Mimo 2.6 on Openrouter, have proved it to be as impressive in practice as it is in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these most recent Flash models is good enough to easily complete most daily tasks.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could likely function as a replacement. Hermes, Deepseek Harness, or Opencode would be my next most likely choices, but I've gotten serious work completed with even the tiniest agents such as NullClaw, ZeroClaw, and PicoClaw. Once a model is given access to files, the OS command line, and an auto-loading file such as agents.md - by any harness - you can get it to complete virtually any required work on your local computer. That provides the basis to prompt it to build a skill system, install and run libraries, write code to create tools of any sort, to control the PC, etc. Codex is most interesting otherwise, for it's build-in computer use and browser control tools, but I haven't needed that so much yet for production development work, and the third party computer use tools work great with it. What matters most is still the raw capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 19Sep 26, 2026 at 14:28

Much of my August 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are to DeepSeek v4.1 Flash (now my daily driver over the APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source models), Meta Muse 1.3 Contributor (the least expensive near-frontier choice), and Qwen 3.8 Flash Next (an amazing option for local use).

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet with that workflow, often running it non-stop for many days in a row - and the models available in ChatGPT always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash models (other bigger models will of course use up your account much more quickly).

GLM 5.3 Flash has been added to Cline-pass, so it's my preferred choice for vision work, and is the first alternate brain I use whenever Deepseek v4.1 has trouble with a task (which is very rare). I've also been totally impressed by the value of Meta Muse 1.3 Contributor. Muse 1.3 is so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely become my preferred choices, because they're currently the best cost per task model rated by Artificial Analysis. My tests of Mimo 2.6 on Openrouter, have proved it to be as impressive in practice as it is in the benchmarks.

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model on the APIs (v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run

It's getting to the point that any of these models is good enough to easily complete most daily tasks.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production development work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM. That is the sweet spot for the class of machines (128Gb shared RAM) that will be relatively affordable for at least the next few years.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. My initial tests of the 2 bit quant of Mimo 2.6 Flash were not successful (it ended up looping during complex tasks). The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b, & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created months ago by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are currently the only possible option, but their viability for production use is questionable: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield some useful productive capability, even on sub-$100 netbooks (that's what I was using in the linked post), but some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught up by any expectations that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 18Sep 26, 2026 at 13:38

Much of my August 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are to DeepSeek v4.1 Flash (now my daily driver over APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source choices), Meta Muse 1.3 Contributor (the least expensive near-frontier model), and the Qwen 3.8 Flash Next model (for local use).

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet, often running it non-stop for many days in a row - and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash models (other bigger models will of course use up your account much more quickly).

GLM 5.3 Flash has been added to Cline-pass, so it's my preferred choice for any vision work, and as an alternate brain whenever Deepseek v4.1 has trouble with a task. I've also been very impressed by the value of Meta Muse 1.3 Contributor. It's so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely be my preferred choices, because they use up less rate limit than other near-frontier models. In my tests of Mimo 2.6 on Openrouter, it's been impressive!

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model (on the API - but v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production development work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are really the only viable option: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield productive capability, even on sub-$100 netbooks (that's what I was using in the linked post!). Some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught in the expectation that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 17Sep 26, 2026 at 13:02

Most of my August 2026 State of Affairs post is still valid: https://aibynick.com/thread/63

The biggest updates are to DeepSeek v4.1 Flash (now my daily driver over APIs), the release of Mimo 2.6 Pro & Flash (now the leading open source choices), Meta Muse 1.3 Contributor (the least expensive near-frontier model), and the Qwen 3.8 Flash Next model (for local use).

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet, often running it non-stop for many days in a row - and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I rely daily on Deepseek V4.1 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DeepSeek Flash models (other bigger models will of course use up your account much more quickly).

GLM 5.3 Flash has been added to Cline-pass, so it's my preferred choice for any vision work, and as an alternate brain whenever Deepseek v4.1 has trouble with a task. I've also been very impressed by the value of Meta Muse 1.3 Contributor. It's so inexpensive on the API, that I've been able to complete entire applications with it, without any usage even registering in the Cline-pass control panel. When Mimo 2.6 Pro and Flash get added to Cline-pass, they will likely be my preferred choices, because they use up less rate limit than other near-frontier models. In my tests of Mimo 2.6 on Openrouter, it's been impressive!

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), Mimo v2.6 flash ($0.14/M input tokens $0.28/M output tokens), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash now supports native vision input, and has superseded the Deepseek v4 Pro model (on the API - but v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting, and Mimo 2.6 Flash appears to be a direct competitor with comparable specs. They both punch very near the quality of frontier models in both capability and knowledge, and are extraordinarily cheap & fast to run.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production development work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash, Mimo 2.6 Flash, and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 3.6 35a3, 3.8 27b & Gemma 4 models.

If you're hardware savvy, you could consider alternately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec specialized PCI connectors, power, and fans, perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they can certainly end up be being a real hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, are still used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. Keep in mind, the following examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, ternary models are really the only viable option: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with ternary models can actually yield productive capability, even on sub-$100 netbooks (that's what I was using in the linked post!). Some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows for me. Just don't get caught in the expectation that ternary models can replace a GPU yet.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 16Sep 26, 2026 at 13:02

Most of what I wrote in the August 2026 Current State of Affairs post, is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet, often running non-stop for many days in a row, and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account much more quickly).

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash has superseded the Deepseek v4 Pro model (on the API - but v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting. It punches very near the quality of frontier models in both capability and knowledge, and it's extraordinarily cheap & fast to run.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 35a3, 27b & Gemma 4 models.

If you're hardware savvy, you could consider altenately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec power, fans, specialized PCI connectors, and to perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they certainly can be be a hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, still appear to come in the form of used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. These examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools. Remember, those apps were all generated 1/2 year ago, on machines that can still be purchased for less than $1000.

For machines without any GPU, try some of the ternary models: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b and smaller models are remarkably fast on pure CPU, for simple text generation, but code produced by the tiny models will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with these ternary models can actually yield productive capability, even on $100 netbooks (that's what I was using the linked post!). Some knowledge and experience goes a long way with those tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows with them.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 15Sep 21, 2026 at 17:59

Most of what I wrote in the August 2026 Current State of Affairs post, is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet, often running non-stop for many days in a row, and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account much more quickly).

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash has superseded the Deepseek v4 Pro model (on the API - but v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting. It punches very near the quality of frontier models in both capability and knowledge, and it's extraordinarily cheap & fast to run.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s.

Using multiple RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 35a3, 27b & Gemma 4 models.

If you're hardware savvy, you could consider altenately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec power, fans, specialized PCI connectors, and to perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they certainly can be be a hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, still appear to come in the form of used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. These examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools.

For machines without any GPU, try some of the ternary models: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running legitimately useful overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b model is remarkably fast on pure CPU, for simple text generation, but code produced by it will likely require revision. Use the 27b version as a final overnight editor/revisor. Learning to work with these models can yield actually productive capability, even on $100 netbooks. Some knowledge and experience goes a long way with these tools. Building basic output with small models and then revising that output with progressively larger models (and/or your own elbow grease), has yielded the most productive workflows with these models.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 14Sep 21, 2026 at 17:56

Most of what I wrote in the August 2026 Current State of Affairs post, is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 . I've never hit a rate limit yet, often running non-stop for many days in a row, and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account much more quickly).

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - a phenomenal buy, but you must agree to share your data publicly), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out). Note that v4.1 Flash has superseded the Deepseek v4 Pro model (on the API - but v4.1 Flash is much too big for most people to run locally. For local self-hosting, v4 Flash is still the Deepseek model to choose).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance, both on the API and for local hosting. It punches very near the quality of frontier models in both capability and knowledge, and it's extraordinarily cheap & fast to run.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work. What matters most is still capability of the LLM(s) you use, and the way you engineer your workflows.

My favorite locally hosted models are:

  • For production implementations, where a 2 DGX Spark cluster is the reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can always add more DGX Spark machines later to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and the comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, far more reliable, and still very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is currently a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed of those machines is faster than Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Using RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less money on hardware, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops can be found intermittently for around $2850 (https://www.amazon.com/gp/product/B0DW238TXK/ref=ox_sc_saved_title_3?smid=A2L77EE7U53NWQ&th=1).

For a rock bottom budget build, look at RTX 5060ti GPUs and the Qwen 35a3, 27b & Gemma 4 models.

If you're hardware savvy, you could consider altenately building a machine with 4 used Telsa V100s (each with 32Gb VRAM), but be prepared to spec power, fans, specialized PCI connectors, and to perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they certainly can be be a hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, still appear to come in the form of used laptops on Ebay, especially those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. These examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools.

For machines without any GPU, try some the ternary models: https://aibynick.com/thread/24?page=1#post-205 . They can actually be successful at running overnight jobs. Note that the earlier Qwen 3.6 27b version seems to be preferrable to the newer 3.8 version, and the 8b version provides a useful mix of capability and speed. The 4b model is remarkably fast on pure CPU, for simple text generation, but code produced by it will likely require revision. Use the 27b version as an overnight supervisor/revisor.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 13Sep 21, 2026 at 17:43

Most of what I wrote in the August 2026 Current State of Affairs post, is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine: https://aibynick.com/thread/3 I've never hit a rate limit yet, often running non-stop for many days in a row, and the models available there always lead the frontier.

For local OS configuration tasks, software installations, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API: https://cline.bot/cline-pass . I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

If you don't want to commit to any subscriptions, on Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - but you must agree to share your data publicly), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out).

GLM 5.3 Flash is currently a genuinely stunning model for price/performance. It punches very near the quality of frontier models, and is cheap & fast.

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a production implementation, where a 2 DGX Spark cluster is a reasonable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality choices. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can add more DGX Spark machines later, as needed, to handle bigger models, or to increase performance.
  • For personal use, on a single DGX Spark, a 128Gb Strix Halo, a machine with an RTX 6000 GPU, or a 128Gb Apple laptop, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash, and Hy3. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below, and a comparison of models building Rubik's Cube demos at https://aibynick.com/thread/54#post-185).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

As a quick demo, here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. 3.8 Flash Next is still a preview model, but it performs fast, is exceptionally capable, and only requires a single machine with 90+Gb VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Using RTX 6000s, you'll get maximum output speed before moving to datacenter class equipment, but you'll still need to spend more than $30,000 to build a system with the same amount of VRAM as 2 DGX Sparks - and you'll likely need to upgrade the electrical circuits in your building. So, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, the DGX Spark (GB10) architecture is hard to beat, for the price and power requirements.

If you need to spend less, consider a single Strix Halo machine running Qwen 3.8 Flash Next (especially the halogen build: https://aibynick.com/thread/79#post-250), a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops still can be found intermittently for around $2850.

For a budget build, look at RTX 5060ti GPUs and the Qwen 35a3, 27b & Gemma 4 models.

If you're hardware savvy, you could consider building a machine with 4 used Telsa V100s (each with 32Gb VRAM), as an alternative, but be prepared to spec power, fans, connectors, and perform lots of software tweaking & deal with failures which require time consuming research. If you've got the knowledge and time, used V100s currently seem to lead in overall price/performance, but they certainly do seem to be a hassle.

The lowest prices I've seen lately for any sort of usable GPU hardware, still appear to come in the form of used laptops on Ebay, such as those with mobile 16Gb 306ti GPUs (around the $1000 price point for an entire portable machine). I've been able to build a lot of absolutely useful software with those machines, especially with the Qwen 3.6 35a3 model. These examples from the Quick Start at https://aibynick.com/thread/29 were all created by Qwen 3.6 35a3:

If you're really budget constrained, that Qwen 3.6 35a3 model is a life saver, with very fast performance. It turns the least expensive GPU hardware into genuinely useful inference tools.

For machines without any GPU, try the ternary models: https://aibynick.com/thread/24?page=1#post-205 They can actually be successful at running overnight jobs. Note that the Qwen 3.6 27b seems to be preferrable to the newer 3.8 ternary version, and the 8Gb version provides a useful mix of capability and speed. Use the 27b version as a supervisor.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 12Sep 21, 2026 at 16:48

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local OS configuration tasks, software installations, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

On Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - but you must agree to share your data), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a production implementation, where a 2 DGX Spark cluster is a usable baseline hardware requirement, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine) are production quality. That hardware, with those models, are appropriate starter configurations for most small-medium business environments where light daily use is anticipated. You can add more DGX Spark machines later, as needed, to handle bigger models, or to increase performance.
  • For personal use, on single DGX Spark, 128Gb Strix Halo, RTX 6000, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months (see the examples below).
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

Here are a couple 1-shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the model to watch, especially when Alibaba releases the fully trained version 4 model on the same architecture. This model performs fast and only requires a single machine with 90+ GB VRAM.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be clustered together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat, for the price and power requirements. If you're hardware savvy, you could consider building a machine with 4 used Telsa V100s (each with 32Gb VRAM), as an alternative, but be prepared to spec power, fans, and connectors, and perform lots of software tweaking. If you need maximum performance, plan on building a system with 2 RTX 6000s (that will cost $30,000+)

If you need to spend less, consider a single Strix Halo machine running Qwen 3.8 Flash Next, a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and/or Step 3.7 flash. The ASUS ROG Flow Z13 Strix Halo laptops still can occasionally be found around $2850.

For a budget build, look at RTX 5060ti GPUs and the Qwen 35a3, 27b & Gemma 4 models.

The best prices I see for usable GPU hardware still seem to be for used laptops on Ebay which have the mobile 16Gb 306ti GPU (around the $1000 price point for an entire portable machine). I've been able to build a lot of quite useful software with those machines, especially with the Qwen 3.6 35a3 model.

For machines without any GPU, try ternary models: https://aibynick.com/thread/24?page=1#post-205 They can actually be successful at running overnight jobs.

For a quick rundown of the most common workflow patterns currently used to build software with LLMs, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 11Sep 21, 2026 at 15:58

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

On Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - but you must agree to share your data), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a 2 DGX Spark cluster, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.
  • For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

Here are a couple 1 shot games generated by locally hosted Qwen 3.8 Flash Next:

Qwen 3.8 Flash Next is really the one to watch, especially when they release the fully trained version 4 model on the same architecture.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat.

If you need to spend less, consider a single Strix Halo machine running Qwen 3.8 Flash Next, a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and Step 3.7 flash.

For a budget build, look at RTX 5060 GPUs and the Qwen 27b, 35a3 & Gemma 4 models.

For machines without any GPU, try the ternary models: https://aibynick.com/thread/24?page=1#post-205

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month: https://aibynick.com/thread/63

Version 10Sep 21, 2026 at 02:32

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

On Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - but you must agree to share your data), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a 2 DGX Spark cluster, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.
  • For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

Here are a couple 1 shot games generated by locally hosted Qwen 3.8 Flash Next:

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat.

If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next, a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and Step 3.7 flash. Qwen 3.8 Flash Next is really the one to watch, especially when they release the fully trained version 4 model on the same architecture.

For a budget build, look at RTX 5060 GPUs and the Qwen 27b, 35a3 & Gemma 4 models.

For machines without any GPU, try the ternary models: https://aibynick.com/thread/24?page=1#post-205

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month:

https://aibynick.com/thread/63

Version 9Sep 21, 2026 at 02:29

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

On Openrouter it's hard to beat the value of GLM 5.3 Flash ($.075/.25 per million tokens in/out), Muse Contributor 1.3 ($.10/.20 per million tokens in/out - but you must agree to share your data), and Deepseek v4.1 Flash ($.15/.52 per million tokens in/out).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a 2 DGX Spark cluster, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.
  • For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat.

If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next, a lower quant of Deepseek V4 Flash, Hy3, Mimo 2.5, and Step 3.7 flash.

For a budget build, look at RTX 5060 GPUs and the Qwen 27b, 35a3 & Gemma 4 models.

For machines without any GPU, try the ternary models: https://aibynick.com/thread/24?page=1#post-205

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month:

https://aibynick.com/thread/63

Version 8Sep 20, 2026 at 02:40

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a 2 DGX Spark cluster, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.
  • For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

2 DGX Spark machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat.

If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next, a lower quant of Deepseek V4 Flash, and Mimo 2.5. For a budget build, look at RTX 5060 GPUs and the Qwen 27b, 35a3 & Gemma 4 models.

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month:

https://aibynick.com/thread/63

Version 7Sep 20, 2026 at 02:31

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

I use Pi as my harness nearly universally, for coding, but for the way I work, nearly any other well known harness could function as a replacement (Hermes, Deepseek Harness, or Opencode would be my next most likely choices). Codex is most interesting otherwise, for computer use, but I haven't needed that so much yet for production work.

My favorite locally hosted models are:

  • For a 2 DGX Spark cluster, GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.
  • For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.
  • For smaller GPUs, I'm not a fan of the ultra popular Qwen 3.8 27b. I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

2 DGX Sparks machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat. If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next or a lower quant of Deepseek V4 Flash. For a budget build, look at RTX 5060 GPUs and the Qwen 27b & 35a3 models.

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month:

https://aibynick.com/thread/63

Version 6Sep 20, 2026 at 02:23

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex is most interesting otherwise, for computer use, but I haven't needed that much yet for production work.

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration tasks, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account more quickly).

My favorite locally hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments that would expect light use. You can add more DGX Spark machines later, as needed.

For single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machines, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.

I'm not a fan of the ultra popular Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable. Perhaps consider Qwen 3.8 27b as an orchestrator model, but have the smaller MOE do as much work as possible.

2 DGX Sparks machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on those machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power, and they support the full CUDA stack, but they don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run models with significant world knowledge and coding capabilities, such as GLM 5.3 Flash and Deepseek V4 Flash, they're hard to beat. If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next or a lower quant of Deepseek V4 Flash. For a budget build, look at RTX 5060 GPUs and the Qwen 27b & 35a3 models.

For a quick rundown of the most common workflow patterns, see the Current State of Affairs article from last month:

https://aibynick.com/thread/63

Version 5Sep 20, 2026 at 02:00

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex is most interesting otherwise, for computer use, but I haven't needed that much yet for production work.

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration task, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account much more quickly).

My favorite locally hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments with light use. You can add more DGX Spark machines later, as needed.

For a single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machine, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.

I'm not a fan of the ultra popular Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable.

2 DGX Sparks machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on these machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power and support the full CUDA stack, but don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run real models like GLM 5.3 Flash, they're hard to beat. If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next, or for a budget build, look at RTX 5060s and 27b or 35a3 qwen models.

For a quick rundown of the most common workflow patterns, see the current state of affairs article from last month:

https://aibynick.com/thread/63

Version 4Sep 20, 2026 at 01:53

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex is most interesting otherwise, for computer use, but I haven't needed that much yet for production work.

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my complicated software development projects with my ChatGPT zip file routine. I've never hit a rate limit yet, often running non-stop for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configuration task, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, using the DS4F model (other bigger models will of course use up your account much more quickly).

My favorite locally hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (the IQ3_XXS quant, because it enables full 1M token context) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models, are what I'd recommend for a production local AI implementation, in most small-medium business environments with light use. You can add more DGX Spark machines later, as needed.

For a single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machine, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic after the next few months.

I'm not a fan of the ultra popular Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model. It's much faster, more reliable, and very capable.

2 DGX Sparks machines plus a QSFP cable can still be purchased for $10,000 or less. The preprocessing speed on these machines is faster than on Mac and Strix Halo machines, and additional units can be linked together to run bigger models, and/or to handle more concurrent users. They require very little power and support the full CUDA stack, but don't have the bandwidth of dedicated RTX 5090s or 6000s. Still, to run real models like GLM 5.3 Flash, they're hard to beat. If you need to spend less, consider a Strix Halo machine running Qwen 3.8 Flash Next, or for a budget build, look at RTX 5060s and 27b or 35a3 qwen models.

Version 3Sep 20, 2026 at 01:52

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex is most interesting otherwise, for computer use, but I haven't needed that much yet for production work.

My 2 best subscription buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my most complicated software development projects with my ChatGPT zip file routine. I've never hit a single rate limit, all year, often running all day for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configs, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, with the DS4F model (other bigger models use up your account very quickly).

My favorite locally hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (IQ3_XXS quant) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models are what I'd recommend for a real starter production local AI implementation, in a small-medium size business environment. You can add more DGX Spark machines later, if needed.

For a single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machine, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic during the next few months.

I'm not a fan of Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model.

Those are currently my most used and trusted AI tools.

Version 2Sep 19, 2026 at 17:37

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex is most interesting to me otherwise.

My best 2 buys are still ChatGPT for $20 per month and Cline-pass for less than $80 per year. That $27-ish per month covers all my real needs.

I still complete most of my most complicated software development projects with my ChatGPT zip file routine. I've never hit a single rate limit, all year, often running all day every, for many days in a row:

https://aibynick.com/thread/3

For local installations, OS configs, and other IT work, I still rely daily on Deepseek V4 Flash via the Cline-pass API. I've never hit a rate limit there either, with the DS4F model (other bigger models use up your account very quickly).

My favorite locally hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (IQ3_XXS quant) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine). That hardware, with those models are what I'd recommend for a real production local AI implementation, in a small-medium size business environment.

For a single DGX Spark, 128Gb Strix Halo, or 128Gb Apple machine, my favorite models are Qwen 3.8 Flash Next, and the 2 bit quant of Deepseek V4 Flash. I expect Qwen 4 will dominate this key hardware demographic during the next few months.

I'm not a fan of Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model.

Those are currently my most used and trusted AI tools.

Version 1Sep 19, 2026 at 17:34

Most of what I wrote in the Current State of Affairs post is still true:

https://aibynick.com/thread/63

I use Pi as my harness nearly universally, for coding. Codex and Opencode are most interesting to me otherwise.

My best 2 buys are still ChatGPT for $20 per month and Cline-pass for a little less than $80 per year. I've been completing most of my most complicated software development projects with my ChatGPT zip file routine, for the entire calendar year. I've never hit a single rate limit:

https://aibynick.com/thread/3

For local installations, OS configs, and other IT work, I still rely daily on Deepseek V4 Flash on the Cline-pass API. I've never hit a rate limit there either, with the DS4F model (other bigger models will use up your account very quickly).

My favorite local hosted models for a 2 DGX Spark cluster are GLM Flash 5.3 (IQ3_XXS quant) and Deepseek v4 Flash (2-4 bit mixed quant running on the Antirez DS4 engine).

On a single DGX Spark or a 128Gb Strix Halo, my favorite models are Qwen 3.8 Flash Next, the 2 bit quant of Deepseek V4 Flash.

I'm not a fan of Qwen 3.8 27b. For smaller GPUs, I think Qwen 3.6 35a3 MOE is still the most practical model.

Those are currently my most used and trusted AI tools.