Of course, this Rubik's Cube comparison is just a tiny cross section of a demo application. This is the sort of work I do every day with LLMs. All I've ever really cared about is end results, because I build software for production, and examples/demos don't matter at all. Unfortunately, most of the real production software never gets shown publicly, because my clients wouldn't want their privately funded bespoke tools, which provide competitive advantage in challenging markets, to be paraded freely in public.
I think the best thing to do though, is to look for genuine comparisons between models. For example:
https://www.youtube.com/watch?v=mzothrIi0cE
That's a fantastic thing about the Internet. There are many thousands of hard working people trying to make a living on platforms like YouTube. They spend their time and life's energy hoping to produce content which is valuable, and much of it truly is. Whenever new models are released, I watch comparisons from a handful of YouTubers. And I read the results at https://artificialanalysis.ai , which compiles the results of benchmark systems created by many other groups. All that together saves me dozens of hours doing rote work myself, putting LLMs through the paces to see real output. Then I'll add my own standardized tests too, to help evaluate features that are most important to the sorts of work I do, and add those comparisons to the pile. Very quickly, that paints a clear picture of what to expect from any new model.