Comparison of self-hosted LLM Rubiks Cube demos

157 views
Nick Antonaccio
Nick AntonaccioAdmin
Sep 27, 2026 at 00:33 (edited, 18 revisions)
#1

This comparison got buried a while back in a topic about Hy3 (https://aibynick.com/thread/54#post-185), so I'm duplicating it here in its own thread, with some new examples and more detailed results.

I performed a telling comparison between Hy3, Deepseek V4 Flash, GLM 5.3 Flash, Qwen 3.8 Flash Next, Qwen 3.6 35a3, Mimo 2.5, and Laguna S2.1, creating a 3D Rubik's cube solver.

These examples were all performed on local hardware. They were all created by the same initial prompt: 'Please create a rubiks_modelname.html 3d rubiks cube that the user can interact with, and which has the option to start with a randomly mixed up cube, and can visually show the steps to solve, as a 3D motion demo'

Qwen 3.8 Flash Next's result was fantastic, but not because it created the nicest UI. It was the only model which actually created a real Rubik's cube solver. In the application it created, you can scramble the cube, or enter any random moves, and the software reasons a complete working solution, from the ground up, using solving logic derived from first principles. All the other applications simply record moves which have been entered in the UI, and play them in reverse. That's a cheat which doesn't include any real logic that can be used to solve actual cubes. The Qwen 3.8 Flash Next model took a long time to complete this solution, but wow, it accomplished everything from a single prompt, basically completely unattended, without any manual iterations (the server timed out a few times so Pi paused, but a simple 'please continue' was all the model required to complete the task). This application was also generated entirely by a single machine (Asus GX10 (DG Spark)), running the IQ4 quant. I'm telling you, this model has been severely underrated by the community. Here's an export of the session:

https://com-pute.com/nick/rubiks_cube_solver--qwen38fnxt--pi-session-2026-09-21T16-08-57-242Z_01a0c4ba-4697-7370-978e-5fd491b6623c.html

I've got to note that Mimo 2.5 also did a great job, very quickly, with some iterations. Mimo's style is more like how a human approaching the problem might be expected to work. It created the simplest possible working solution, and then I worked with it in steps, to add features. It tends to break down problems into engineering steps, and confidently completes reasonable improvements with interactive guidance. That's a style which feels to me very controllable, with more human intention involved. Mimo 2.6 flash has now been released on Openrouter - I'm very excited to try its open weights. This is another open source model family which doesn't get enough attention. Here's the session export:

https://com-pute.com/nick/rubiksmimo25--pi-session-2026-08-09T18-44-07-981Z_019fe7d6-e4ad-79ff-a6b9-0b5d65bbac81.html

As always, the output from Qwen 3.6 35a3 is utterly impressive for the size and speed of that model. That MOE LLM runs very quickly on laptops which I've purchased for as little as $800, and which have only 16GB of VRAM. Those sorts of used machines with mobile RTX 3080ti GPUs are still found regularly on Ebay for around $1000, and the Qwen models make them actually capable of achieving real software development tasks. The result above was from a single prompt, without any iterations. That's truly a fantastic outcome for such a small model.

GLM 5.3 Flash and Deepseek V4 Flash have been my most used, most trusted self-hosted models, but for this task, Hy3 was most efficient. It got the job done surprisingly quickly, right out of the gate, with the fewest iterations and issues, and built a nice UI. Hy3 only requires a single machine too, where the GLM 5.3 Flash result was created with a 2 DGX Spark cluster. Here's an Hy3 Pi session export for one of the Hy3 generations - note that this session was from a full precision version of Hy3 hosted on Openrouter. I no longer have the session export which used the locally hosted version - all other sessions in this case study were from locally hosted model generations:

https://com-pute.com/nick/rubikshy3--pi-session-2026-08-09T13-27-05-120Z_019fe6b4-a0a0-7991-94e5-3bf74198f611.html

Here is the session export for the Deepseek V4 Flash Antirez DS4 mixed 2-4 bit quant, running on a single DGX Spark. The 4 bit quant which requires 2 clustered DGX Spark machines is actually significantly more capable and reliable, so this result is more impressive than it appears at first. The 2-bit model worked fine, but did require some feedback:

https://com-pute.com/nick/rubiks_cube_local_ds4f--pi-session-2026-08-08T14-41-22-072Z_019fe1d2-4698-7140-aaeb-cb53cdda567b.html

Here's the export for the GLM 5.3 Flash session, which used the IQ3_XXS quant on 2 clustered DGX Spark machines. It didn't require any iterations:

https://com-pute.com/nick/rubiks-cube--glm5.3-flash--pi-session-2026-08-30T14-25-45-086Z_01a0530f-e27e-7c43-bc40-9c30a4003386.html

And here is the Laguna S2.1 session export - it was a disappointing failure compared to all the others:

https://com-pute.com/nick/rubiks-cube-lagunaS21-partial-fail--pi-session-2026-08-08T22-03-31-622Z_019fe367-15a6-72c7-8330-11a0cac3207f.html

Here's a version created by Muse Spark 1.3 Contributor, on the API (not on a locally hosted machine):

https://com-pute.com/nick/rubiks_modelname--musespark13contributor.html

Nick Antonaccio
Nick AntonaccioAdmin
Oct 06, 2026 at 12:18 (edited, 1 revision)
#2

Of course, this Rubik's Cube comparison is just a tiny cross section of a demo application. This is the sort of work I do every day with LLMs. All I've ever really cared about is end results, because I build software for production, and examples/demos don't matter at all. Unfortunately, most of the real production software never gets shown publicly, because my clients wouldn't want their privately funded bespoke tools, which provide competitive advantage in challenging markets, to be paraded freely in public.

I think the best thing to do though, is to look for genuine comparisons between models. For example:

https://www.youtube.com/watch?v=mzothrIi0cE

That's a fantastic thing about the Internet. There are many thousands of hard working people trying to make a living on platforms like YouTube. They spend their time and life's energy hoping to produce content which is valuable, and much of it truly is. Whenever new models are released, I watch comparisons from a handful of YouTubers. And I read the results at https://artificialanalysis.ai , which compiles the results of benchmark systems created by many other groups. All that together saves me dozens of hours doing rote work myself, putting LLMs through the paces to see real output. Then I'll add my own standardized tests too, to help evaluate features that are most important to the sorts of work I do, and add those comparisons to the pile. Very quickly, that paints a clear picture of what to expect from any new model.

Please login to post a reply.

© 2026 AI By Nick.