Post History

Current version by Nick Antonaccio

Current VersionSep 04, 2026 at 12:02

The pace of LLM improvement is dizzying.

To put things into perspective, here's a little story from yesterday:

I used a $100 credit to build a Flask application builder tool. This software development project was an attempt to replace Anvil, Baserow, and Jam.py frameworks, baking some of their best no-code database and UI features/approaches into one tool which outputs Flask code - instead of building another proprietary framework to support the builder features. The goal of the project is to build a deterministic, lightweight, no-code Flask development tool which is competitive with the productivity of LLM based software development practices, while providing applications that are compatible with my familiar LLM development workflows.

This project was not a lightweight challenge, but Fable got it done in a single casual day, just taking a bit of my spare time occasionally every few hours.

That project ended up being called Flaskforge, and it was the sort of thing that would have taken months of dedicated professional human labor to create. I've been developing software for about 40 years, and this would have been a daunting, long term project to take on. Working on It would have taken up a huge portion of my life.

Here's the thing: I used Fable 5 on this project because the $100 credit was limited to that model (as opposed to the newer 5.1 release of Fable). What really struck me about how this project evolved was that in the planning stages, I had GPT 5.6 and Fable go back and forth, evaluating each other's plans and implementations, and it was clear that GPT 5.6 bested Fable's work repeatedly. GPT 5.6 on medium caught many errors by Fable, and Fable regularly acknowledged that GPT 5.6 engineered and executed better solutions.

Now GPT 6 (Astra) is being rolled out, and it's far better than both GPT 5.6 and Fable 5.1. Fable, the monstrously powerful LLM which changed cybersecurity forever, and which was able to build a project that would have taken me many months, perhaps years, of hard work to engineer and build well, is now old news.

That's hard to imagine, but progress is clearly continuing. This new OpenAI 'Astra' model, is the one which broke into HuggingFace during internal testing phases at OpenAI. The technical and conceptual challenges it surmounted in that breech were astounding. It's also the model which solved 10 long standing math problems that none of the best human mathematicians have been able to solve for many years. And by all the benchmark results, and endless initial reports by human beta testers, Astra is better at coding, science, and many other critical capabilities that will push technology forward, than any other model which has existed yet. Greg Brockman called this model the beginning of the AGI era.

Astra has apparently made the run-time thinking process far more efficient and effective, by using a looping method that generates multiple tokens at once, without requiring as much compute. I'm waiting to hear more details about exactly how this has been implemented, but we can be sure that other LLMs will include these sorts of approaches very soon...

So the world of artificial intelligence is continuing to change, the rate of that change is increasing dramatically, and it doesn't show any signs of stopping.

I can't help reflect: In September of 2022, several months before ChatGPT was released, I created a little Anvil app which connected with the OpenAI API to query text-babbage-001, text-curie-001, and Davinci models. Those LLMs were 1.3, 6.7, and 175 billion parameters - they could hardly do anything useful. At the time, though, it was just amazing that a non-deterministic machine could form sentences, and actually show some signs of intelligent, meaningful content generation.

Fast forward to just last year, and we were using GPT-4 class models, which were still relatively simple tools with many shortcomings, but they had dramatically changed how much work I was able to accomplish in a day.

Now we're using tools which far outperform many creative and technical human intellectual capabilities, and which operate at 1 million times the speed of human thought.

Astra's capabilities are real, right now, and they are about to be deployed to all of humanity. I can't help but extrapolate about the trends we've seen develop in the short time LLMs have existed. It's really hard to imagine where we will be next year, and I don't expect that we can even begin to comprehend what the capabilities of AI systems might be in 10 years, especially when that unfathomable intelligence will likely be walking around in billions of equally physically capable humanoid robots.

What a time to be alive.

Previous Versions
Version 3Sep 04, 2026 at 12:02

The pace of LLM improvement is dizzying.

I used a $100 credit yesterday to build a Flask application builder tool. This software development project was an attempt to replace Anvil, Baserow, and Jam.py frameworks, baking some of their best no-code database and UI features/approaches into one tool which outputs Flask code - instead of building another proprietary framework to support the builder features. The goal is to build a deterministic, lightweight, no-code Flask development tool which is competitive with the productivity of LLM based software development practices, while providing applications that are compatible with my familiar LLM development workflows.

This project was not a lightweight challenge, but Fable got it done in a single casual day, just taking a bit of my spare time occasionally every few hours.

That project ended up being called Flaskforge, and it was the sort of thing that would have taken months of dedicated professional human labor to create.

Here's the thing: I used Fable 5 on this project because the $100 credit was for that model. What really got me about how this project evolved was that in the planning stages, I had GPT 5.6 and Fable go back and forth, evaluating each other's plans and implementations, and it was clear that GPT 5.6 bested Fable repeatedly. GPT 5.6 on medium caught many errors by Fable, and Fable regularly acknowledged that GPT 5.6 engineered and executed better solutions.

Now GPT 6 (Astra) is being rolled out, and it's far better than both GPT 5.6 and Fable 5.1. That's hard to imagine, but progress is clearly continuing. This model, 'Astra', is the one which broke into HuggingFace during internal testing phases at OpenAI. The technical challenges it surmounted in that breech were astounding. This is also the model that solved 10 long standing math problems which none of the best human mathematicians have been able to solve for many years. And by all the benchmark results, and all initial reports by human beta testers, Astra is better at coding, science, and many other critical capabilities than any other model which has existed yet. Greg Brockman called this model the beginning of the AGI era.

Astra has apparently made the run-time thinking process far more efficient and effective, by using a looping method to generate multiple tokens at once. I'm waiting to hear more about how this works.

The world of artificial intelligence is continuing to change, and the rate of that change is improving dramatically. In September of 2022, before ChatGPT was released, I created a little Anvil app which connected to the OpenAI API to query text-babbage-001, text-curie-001, and davinci. Those models were 1.3, 6.7, and 175 billion parameters - and they could hardly do anything useful. It was just amazing that they could form sentences, and actually show some signs of intelligent, meaningful generation. Fast forward to just last year, and we were using GPT-4 class models, which were still relatively simple tools with many shortcomings. Now we're using tools which far outperform many creative and technical human intellectual capabilities, and which operate at 1 million times the speed of human thought.

Astra's capabilities are real, right now, and are about to be deployed to all of humanity. I can't help but extrapolate about the trends we've seen develop in the short time LLMs have existed. It's really hard to imagine where we will be next year, and I don't expect that we even comprehend what the capabilities of AI might be in 10 years, when unfathomable intelligence will likely be walking around in billions of equally physically capable humanoid robots. What a time to be alive.

Version 2Sep 04, 2026 at 11:32

The pace of improvement is dizzying.

I used a $100 credit yesterday to build a Flask application builder tool. This project was an attempt to replace Anvil, Baserow, and Jam.py frameworks, baking some of their best no-code database and UI features/approaches into one tool which outputs Flask code - instead of building another proprietary framework to support the builder features.

This was not a lightweight challenge, but Fable got it done in a single day, just taking a bit of my spare time.

That project ended up being called Flaskforge, and it was the sort of thing that would have taken months of dedicated professional human labor to create.

Here's the thing: I used Fable 5 on this project because the $100 credit was for that model. What really got me about how this project evolved was that in the planning stages, I had GPT 5.6 and Fable go back and forth, evaluating each other's plans and implementations, and it was clear that GPT 5.6 bested Fable repeatedly. GPT 5.6 on medium caught many errors by Fable, and Fable regularly acknowledged that GPT 5.6 engineered and executed better solutions.

Now GPT 6 (Astra) is being rolled out, and it's far better than both GPT 5.6 and Fable 5.1. That's hard to imagine, but progress is clearly continuing. This model, Astra, is the one which broke into HuggingFace during internal testing phases at OpenAI. he technical challenges it surmounted in that breech were astounding. This is the model that solved 10 long standing math problems which none of the best human mathematicians have been able to solve for many years. And by all the benchmark results, and all initial reports by human beta testers, Astra is better at coding, science, and many other critical capabilities than any other model which has existed yet.

The world of artificial intelligence is continuing to change, and the rate of that change is improving dramatically. Last year we were using GPT-4 class models which were useful tools. Now we're using tools which best many human intellectual capabilities, and which operate at 1 million times the speed of human thought. This capability is real, right now, and about to be available to all of humanity. Where will we be next year? Can we even imagine where we'll be in 10 years, when unfathomable intelligence is walking around in billions of equally physically capable humanoid robots?

Version 1Sep 04, 2026 at 08:20

The pace of improvement is dizzying.

I used a $100 credit yesterday to build a Flask builder application. This project was an attempt to replace Anvil, Baserow, and Jam.py frameworks, taking some of their best no-code database and UI features into one tool which outputs Flask code, instead of building another proprietary framework to support the features. This was not a lightweight challenge, but Fable got it done in a single day, just taking a bit of my spare time.

That project ended up being called Flaskforge, and it was the sort of thing that would have taken months of dedicated professional human labor to create.

Here's the thing: I used Fable 5 on this project because the $100 credit was for that model. What really got me about how this project evolved was that in the planning stages, I had GPT 5.6 and Fable go back and forth, evaluating each other's plans and implementations, and it was clear that GPT 5.6 bested Fable repeatedly. It caught many errors by Fable, and Fable regularly acknowledges that GPT 5.6 engineered and executed better solutions.

Now GPT 6 (Astra) is being rolled out, and it's far better than both GPT 5.6 and Fable. That's hard to imagine, but the progress continues. Astra is the model that broke into HuggingFace during testing (the technical challenges it surmounted in that breech were astounding). This is the model that solved 10 long standing math problems which none of the best human mathematicians have been able to solve for many years. And by all the benchmark results, and all the initial reports, Astra is better at coding, science, and other critical capabilities than any other model which has existed yet.

The world is continuing to change, and the rate of that change is improving dramatically. Last year we were using GPT-4 class models which were useful tools. Now we're using tools which best many human intellectual capabilities, at 1 million times the speed. This capability is real, right now, and about to be available to all of humanity. Where will we be next year? Can we even imagine where we'll be in 10 years, when an unfathomable intelligence is walking around in billions of equally physically capable humanoid robots?