This trend doesn't necessarily mean that all the current ecosystems are dead or dying. I still use LM Studio to run and test a huge variety of models on all my servers. And in many cases, I don't need better optimization. Qwen 3.6 35a3 runs at 60+ tokens per second on the DGX Sparks and Halo machines, even at 8 bit quantization - and it's as good now at completing the tasks I have come to trust it with, as the day it first blew me away. So that model doesn't need optimization in LM Studio. It works fine for now, out of the box. And Qwen Flash Next runs at 25-ish tokens per second on a single Strix-Halo machine, in LM Studio, on Windows OS, without any optimizations beyond selecting MTP with a preview depth of 3.
That's super useful, without having to dual-boot Linux, install Halogen, and dedicate that machine to basically being a Halogen-only box (until the next better platform emerges...).
The thing is, more and more, this narrow optimization path is actually turning out to be the best solution, and it looks like we'll more and more often end up needing to run dedicated boxes like that, outside of testing new releases.
For example, I've currently got a cluster of 2 DGX Sparks which are basically always just running GLM 5.3 Flash, and another cluster running Deepseek V4 Flash in Antirez DS4. There's rarely a reason to shut down those models, because they're my work horses. For most users with an RTX 5090, it currently makes sense to just run Qwen 3.8 27b in NInfer all the time, because you're going to be hard pressed to find a better, smarter, faster performing production model + engine, overall, for that hardware.
I may concurrently install a small specialized model on one of the machines in my GLM 5.3 Flash cluster, for example to generate music, but for the most part I just want those machines to run GLM 5.3 Flash. I'll use my older machines with 16Gb VRAM to run Qwen 3.6 35a3, for example, which makes them still very useful hardware - but then again, I won't run much else on that hardware, because that model is basically the only one which is really worth running on those machines.