Kimi K3

251 views Pinned
Nick Antonaccio
Nick AntonaccioAdmin
Jul 19, 2026 at 20:21 (edited, 2 revisions)
#1

The news about Kimi K3 is everywhere, so I'll keep my notes about it short.

The general consensus seems to be that Kimi K3 is genuinely disruptive as a frontier class model. Moonshot's own published stance is that K3 doesn't beat Fable and GPT 5.6 Sol in many ways, but at the same time, K3 is ranked #1 in the Front End Code Arena, in blind tests by the community at arena.ai.

My own experience with it seems to indicate that K3 is a legitimate contender right at the forefront of frontier capability, but like everything else in AI, there are jagged edges.

I think, in general, Kimi K3 is certainly much more capable than any other open source model, including even GLM 5.2. From all the published community demos, Kimi K3's coding capabilities are absolutely right up there with Fable and GPT 5.6 Sol, especially when it comes to building 3D apps, front end code, and in some other classes of work.

Many demonstrated 3D game vibe coding comparisons, and web site graphics demos involving 3D, show Kimi K3 producing even more impressive first-shot results than Fable and GPT 5.6 Sol. And K3 is capable of regularly running for hours to complete very complex code base tasks.

What I can add to all this is that Kimi K3 produced the most impressive results in the world-knowledge tests I've been running for years on LLMs, which include a suite of questions about the topics of paramotor and obscure programming libraries (edit: also be sure to check out the impressive R3 project case study below!)

Fable clearly has more depth of knowledge about obscure topics. It produced dozens of pages of detailed, correct information about my own involvement in the paramotor industry, including all the details contained in hundreds of pages of instructional texts I've written (which are established in a niche, but certainly not mainstream knowledge). Fable also produced dozens of pages about texts I've written about programming with Rebol, Anvil, and other languages/frameworks.

Fable a 10 trillion parameter model, so it has more knowledge than any other model, including the less than 3 trillion parameter Kimi K3. Those sorts of successful test results about obscure knowledge, therefore, are expected.

To be clear, all those knowledge tests I performed, are always done in a way that restricts the model from performing any sort of web search, or using any tools that enable the model to research any information available online. They test only the depth, details, and correctness of information stored in the models' trained parameters. You can read a bit about that in this post:

https://aibynick.com/thread/45#post-112

The thing that struck me about my comparisons between Fable and K3, were that the general answers I got from K3, about paramotor topics, were actually more useful than what Fable provided initially.

Kimi K3 knew nothing about my writings or my involvement in paramotoring - Fable clearly knows more of the total information collected by all of humanity, especially when it comes to obscure information. But, here's the thing: knowledge of every possible obscure fact ever collected by human activity isn't necessarily required to make a model useful for the most common and practical tasks, which affect daily human operations.

The useful information I got from K3 would have satisfied research being performed about the topic, by someone who knew nothing about paramotoring - better then the results I got from Fable. K3 didn't know obscure details about the topic, but it provided better top level results.

I expect Kimi K3 certainly has enough general knowledge about the world to help it reason solidly about virtually any topic/task that you might ask it to perform. For example, when building software for a client, it's important for a model to understand how typical business processes function. K3 certainly has enough knowledge to understand how most common business processes operate, and it should be able to connect the dots about how most software solutions would be expected to work in common situations.

And of course, every reasonably sized model can perform research online and reason, in context, about obscure details of niche topics. Even models like Deepseek V4 Flash accomplish that sort of work fantastically well - and models of that size don't require datacenter class hardware and electricity to run.

So I think the big takeaway from comparisons about huge frontier models, is that much of the time, most of their deep capabilities just aren't needed, but when they are, the goal is to focus on when/how/why the deep obscure knowledge available in massive models is actually useful.

For most daily work, Deepseek V4 Flash actually has enough capability and built-in knowledge to research virtually any software project I may want to compile and install in normal work. It knows enough about any common OS (Windows 11, Unbuntu 24.04, Android, etc.) to perform typical system management work - and anything it doesn't know, it can look up online, and it can reason about how to use that information to complete a task. V4 Flash can build just about any application functionality needed for typical business software, in common programming languages/frameworks such as Python/Flask, HTML/CSS/JS, etc., and it can even build interesting 3D games, and perform significantly deep research about virtually any topic, if you provide it tools to search the Internet. I'll post the differences between what V4 flash knew about a particular class of paramotoring topics, compared to what Kimi K3 and Fable knew, in a separate post below.

Models with several hundred billion parameters, such as Deepseek V4 Flash, Mimo 2.5, Stepfun 3.7, Minimax, etc. can be run on local hardware, and they can be run fast on LLM API provider services. You can run V4 Flash all day every day for just a few cents per million tokens. It's a competent daily driver for a huge percentage of work that needs to be accomplished by LLMs.

Models like Fable and Kimi K3 are useful when you need to perform more specialized tasks which require lots of built-in knowledge. For example, over the past few days, I've converted my old Merchants' Village software from 2010, written in the Rebol programming language, to Python Flask.

That task required the model to have a deep understanding of the Rebol programming language, in order to make that language conversion possible. Most smaller models don't know anything about Rebol. The smaller models' world knowledge is just too limited. They're trained on popular languages/frameworks that are most likely to be useful in mainstream daily work, but there's just not enough room in the parameters of a 200-300B model, to train it to be an expert on an obscure mostly now-dead language like Rebol. And becoming an expert in Rebol is too big a task to perform with in-context research/learning, even for models with a million token context window.

So for tasks like that, a model like Fable is a fantastically useful tool. It knows everything about everything, natively, without having to research anything online. It's an expert in even the most obscure topics, without ever having to look up anything. That's important when performing complex reasoning, and completing complex development tasks that involve obscure expertise.

And I think K3 sits at a very useful spot in that specialization continuum. It seems to know a lot about all the most useful specializations - things like 3D application development, which smaller models aren't so great at. It may not know everything about obscure tools like Rebol, but who's going to need that sort of specialized, deep, obscure knowledge, regularly?

So much of human work doesn't involve expertise in the most obscure topics. Most useful development work can be accomplished using mainstream languages and well-known tools.

If a model like Kimi K3 can outperform Fable on a particular task, such as building a particular class of 3D application, then it's more useful than Fable, especially if it costs much less, uses much less energy, can potentially be self-hosted, etc., to complete that same class of more commonly useful tasks.

And the same goes for much smaller models, too, by the way. Deepseek V4 Flash is capable enough to complete the overwhelming variety of IT configuration, installation and development work I see every day, and it's much faster, virtually free in comparison - and I can self host it on relatively inexpensive hardware, so I can set it up for clients, to run HIPAA compliant data management tasks, to process private information, to perform endless agentic work for only the cost of the electricity required, etc.

The big thing to watch is how this affects the industry. The release of Kimi K3 is definitely another Deepseek moment, which will affect the balance of power on a global scale. I expect Chinese models to dominate the AI race over the next year, even if they don't ever fully beat the best which Anthropic and OpenAI have to offer. The truth is, most people just won't regularly need the capabilities of those gigantic models - not even professional software developers, doctors, engineers in other fields, content creators, etc.

The race to produce the most capable models will continue, and upcoming super-human capable models will benefit humanity. I'm so glad I was able to use Fable to convert my old Merchants' Village software from Rebol to Flask. Having that sort of capability is clearly useful, and frontier class models will continue to stretch the capabilities of what humans can achieve, so it's worthwhile to continue that research and development work, in my opinion.

But I think we're beginning to see the end of what companies like Anthropic and OpenAI will be able to achieve. They're going to run out of money, and when (not if) they fail, the US and the world will likely face serious financial repercussions. Too much of our economy is invested in the development of frontier models. There is no way for those companies to be financially viable in the end. Their goals require more financial resources than can be sustained by demand. Smaller models will win in the economy, especially as those smaller models get to be even more capable than they are now, which is inevitable.

At the moment, it looks to me like China is poised to win the long-term frontier model battle, because they're using less powerful hardware to build models which, as of Kimi K3, compete right at the front lines with the US's best model capabilities - and because the Chinese model of government and economy better suits the success of this sort of large scale development effort: as opposed to every man/company for themselves, competing to win the war on their own (the US way), every resource in their economy instead works competitively to make the country successful as a whole. China may not produce a single trillionaire, but the resources of their entire economy are more likely to be sorted out as a single system, to produce the best outcome for their entire country.

We're seeing the Chinese system work, in the release of Kimi K3, and I expect that pattern to continue, while many American companies are likely to eventually die fighting each other, if they don't pool their resources and join together, somehow. Perhaps an American monopoly will arise, as one company pulls ahead and the others die off. I'm not sure how that may play out, but I certainly don't expect that the US industry will be able to continue the way it has for much longer (circular investments, growing debt at a faster rate than income, etc.).

This is all such an amazing turn in the course of human history. In retrospect, I think the release of Kimi K3 will be seen as a significant point in history - an open source model which is more intellectually capable, in many jagged ways, than most humans.

I'm looking forward to seeing all the smaller models distilled from K3, which we'll be able to run on self-hosted hardware, and I'm looking forward to seeing all the other companies employ the novel improvements to LLM technology which Moonshot has given away with this magnificent release. The general improvements to AI as a whole will continue to be exciting going forward, and it's awesome to see how far we've come already, just in a few years.

I've said it 1001 times: humanity is just taking the first steps in a million mile journey with artificial intelligence. It's incredible to watch, and it's going to be even more fantastic to watch in the coming years.

Nick Antonaccio
Nick AntonaccioAdmin
Jul 18, 2026 at 17:05
#2

I posted comparisons of model output here:

https://aibynick.com/thread/45?page=1#post-158

Nick Antonaccio
Nick AntonaccioAdmin
Jul 19, 2026 at 16:59 (edited, 4 revisions)
#3

This is amazing example of a project I would have paid a lot of money to get completed a decade ago:

http://xqx1.com:8080

That project is a rebuild of the old Rebol 3 programming language interpreter (R3), with its 'View' GUI dialect, networking, file handing, graphics, and other features, for WASM, so that it runs entirely in a web browser

You can paste this line into the console at the link above, to see some little example apps run:

do http://com-pute.com/nick/demo.r3

That rebuilt interpreter was a tag team effort between Fable and Kimi K3, completed in a day. Deepseek V4 Pro did some good work during the process, but it couldn't keep up with the power of Fable and K3.

My girlfriend had a day of rate-limited Fable usage left on her $20/month Claude account, which would have otherwise gone to waste, so I uploaded the source code to my old Android version of R3 with View, and let it cook on building a version that runs entirely in the browser. What Fable produced, after hitting the rate limit a few times, was a language interpreter that mostly worked - but it was missing a lot of functionality.

I asked Fable for a source code package and a prompt to paste into Pi on one of my VPS accounts, to get it compiled and working. First shot, the interpreter generally worked in the browser, but none of the GUI examples ran.

So I let Deepseek V4 Pro hack at it for a while, and it actually did get some GUIs to display, but Deepseek just continually ran into lots of issues. So I switched Pi over to Kimi K3 before going to bed - and woke up to a functioning version of Rebol 3 with View running in the browser. This was a seriously impressive result.

I tested a few scripts from https://learnrebol.com/rebol3_book.html and quickly realized that networking calls (reading of data at HTTP URLs) weren't functioning correctly. Kimi K3 took less than a minute to get that worked out. Then I noticed alerts in the UI weren't working. K3 got them going with a prompt.

The most impressive thing about this build was how deeply Fable and Kimi K3 understood everything about R3 code, not just the compiled source code, but everything about the deep roots of the REBOL language, all the way up through not only Rebol code/data structures, but also dialects.

These models are not trained on much about Rebol, so there was an enormous amount of genuinely powerful reasoning and in-context learning going on during this process. Deepseek did an impressive job for what it costs, but it struggled for a while, and I got tired of waiting during Fable rate limits, so just sicked Kimi K3 on the remaining issues. Beyond the rate-limited use of Fable, which would have gone unused, the total cost to get this project up and running to the point where it stands now, was $3.88 - mostly for the Kimi K3 tokens on OpenRouter.

It was awesome to watch Fable and Kimi K3 work - they're both absolute beasts.

There are many more details about this project, along with all the downloadable files, session transcripts, etc., here:

https://rebolforum.com/topic/834

Nick Antonaccio
Nick AntonaccioAdmin
Jul 21, 2026 at 06:16 (edited, 2 revisions)
#4

Here's Kimi K3's entry into the 'provide an html 3D driving game' comparison:

https://com-pute.com/nick/driving-game--kimik3c.html

Compared to all the other quick 3D driving games, which I've had every other LLM create so far, this one is the best. It's the only one yet with curved roads that rise and fall.

For comparison, there are links to many of the other 3D driving demo game demos created by other LLMs, in the quick start tutorial at https://aibynick.com/thread/29

Nick Antonaccio
Nick AntonaccioAdmin
Jul 23, 2026 at 13:50 (edited, 10 revisions)
#5

And here are a few little 3D demos using the R3 draw dialect (runnable in the WASM interpreter, in the case study above, at http://xqx1.com:8080 ), created by Kimi K3:

do http://com-pute.com/nick/3d_demo_oneline.r3

do http://com-pute.com/nick/3d-torus-oneline.r3

do http://com-pute.com/nick/3d-blaster-oneline.r3

Again this is impressive because most LLMs don't really know how to write R3 code at all. Kimi K3 learned how to do this in-context, within a single session in Pi.

After making the first few 3D demos, I had Kimi create a skill, to enable any LLM model to create 3D apps for this version of R3:

https://com-pute.com/nick/r3-wasm-3d-skill-3b.zip

These little games were created by lowly Deepseek V4 Flash, using that skill, in Pi:

do http://com-pute.com/nick/3d-driver-oneline.r3

do http://com-pute.com/nick/arena-watch-oneline.r3

Please login to post a reply.

© 2026 AI By Nick.