Post History

Current version by Nick Antonaccio

Current VersionSep 03, 2026 at 12:44

This build cost $700 at the end of August 2026:

https://www.youtube.com/watch?v=KoMtZWyy9Wo

In another video, he demonstrates a few models running on the hardware:

https://www.youtube.com/watch?v=8Hcbr95dkGY

He runs the Unsloth Q4_K_S version of Qwen 3.8 27b with only a 4096 context length, at 3.96 tokens per second. There are definitely better quants of this model that would make better use of this class of hardware (the Ridge quant by Empero AI is one).

So although not ideal, a system like this can actually be viable for some sorts of work. As models improve, we'll start to see cheap hardware become more capable.

Previous Versions
Version 1Sep 03, 2026 at 12:44

This build cost $700 at the end of August 2026:

https://www.youtube.com/watch?v=KoMtZWyy9Wo

In another video, he demonstrates few models running on the hardware:

https://www.youtube.com/watch?v=8Hcbr95dkGY

He runs the Unsloth Q4_K_S version of Qwen 3.8 27b with only a 4096 context length, at 3.96 tokens per second. There are definitely better quants of this model that would make better use of this class of hardware (the Ridge quant by Empero AI is one).

So although not ideal, a system like this can actually be viable for some sorts of work. As models improve, we'll start to see cheap hardware become more capable.