Looking at the refutations of Zitrons predictions in TFA, it boils down to two categories:
1. Zitron claims model capability has peaked
2. Zitron claims AI lab growth (user and revenue) has stalled.
In the first case, TFA refutes by claiming 'wrong' repeatedly, which does not convince me of anything. If anything, Zitron is probably right in this regard, since the majority of progress in recent LLM tech has been setting up of guardrails to cajole the models using 'agents'.
In the second case, numbers are given showing growth, directly refuting Zitron. However, I'm giving Zitron the benefit of the doubt, given the old saying - market can remain irrational far longer than you can remain solvent. As long as people can be convinced that the sky is falling, rational predictions rarely pan out.
> 1. Zitron claims model capability has peaked … Zitron is probably right in this regard …
Is there any objective measure that shows this?
Would we use examples such as his Feb 2024 claim "I believe we're reaching the upper limits about what generative AI can do"?
> 2. Zitron claims AI lab growth (user and revenue) has stalled.… I'm giving Zitron the benefit of the doubt …
Is there some date by which you'd say it'd be fair to evaluate whether Z’s claims are true (without the benefit of the doubt)?
You mentioned revenue - would we use claims such as his 2024 claim that the companies no longer knew how to grow? But that in 2024, 2025, and 2026 both the companies revenues and profits have grown at double-digit rates each period?
You also mentioned users - would we use claims that "Sundar Pichai wants Gemini to be 'used by 500 million people before the end of 2025, 'a number so unrealistic that someone at Google should have been fired, and that someone is Sundar Pichai", where Gemini then hit 750 M users?
Or by what measures should we evaluate whether Z's claims are true?
I'll paste my response to your pre-edited questions:
I think atomic weapons peaked during the trinity test. The subsequent creation of bigger explosions by packing more fissile material does not meaningfully improve the technology.
Since LLMs are a generative technology, let's use a generative skill - painting. There are tons of brilliant painters, all with their own styles. What I think is consistent among master painters is their effortlessness in their craft. Honing a skill makes subsequent attempts less effortful.
To me, LLMs are more like atomic bombs than master artisans. To get better results, add neural nets. What would be meaningful to me, is to constrain models to a fixed amount of compute, run them on any of the copious amount of benchmarks, and see if they get the same scores at increasingly faster speeds. This, at least to me, signals mastery.
On the date for evaluation, I will set my own prediction instead, that AI labs will not become profitable, local models will drain their moat.
Taken on its own, I will concede that Zitron's predictions are wrong, but it is exactly why I bring up market irrationality. The multiple rounds of funding is propping up the unsustainable business model. Without it, user numbers and revenue can't grow.
I forsee real innovation in the AI space after the bubble pops.
Just this morning, I had to read through an LLM response about how a PC8-M5 fitting has a high flow rate because it connects to an 8mm tube, completely ignoring that the threaded M5 on the other end will only have space for a 2mm hole, so it still seems quite similar to the LLMs of 2020.
> In the first case, TFA refutes by claiming 'wrong' repeatedly, which does not convince me of anything. If anything, Zitron is probably right in this regard, since the majority of progress in recent LLM tech has been setting up of guardrails to cajole the models using 'agents'.
You don't have to like them, use them, or consider them "good enough", but the idea that models haven't gotten better in the last two years is ridiculous.
> In the first case, TFA refutes by claiming 'wrong' repeatedly, which does not convince me of anything. If anything, Zitron is probably right in this regard, since the majority of progress in recent LLM tech has been setting up of guardrails to cajole the models using 'agents'.
I am certain that plugging a circa 2023 model into a 2026 harness would be a pretty frustrating experience. Yes, you could code a bit with AI in 2023, but models are just much better at it than they used to be. And smaller open models are leaps and bounds better at it than they were three years ago.
Agreed, though I do think that LLMs are still more similar than we think. Sometime after the first release, AI labs found that coding sat in the niche space of lots of easily digestible data and fast feedback from error messages and compiler checks etc. This allowed models to be trained with a focus on coding tasks, but the underlying technology is still the same, the infrastructure around it changed, they are still generating via probabilistic sampling.
Don't get me wrong, I'm using local coding agents myself with varying levels of success and frustration, but the models themselves behave similarly to their siblings from 2020.
The infrastructure improved, that includes the data. I have a pet theory that if they took the earlier models and retrain it with the data they used to train the latest models, we will get a similar result.
This is kind of a "I'm not here to tell you about Jesus he either lives in your heart or he doesn't".
I am an engineer, I have about 15 years of experience. I have been using AI in my job since mid 2025. Over that time it has gone from being an interesting toy that could kind of help but would often hinder, to being an absolutely explosively powerful tool. Just from personal experience it is the thing that has improved the most of any of my tools in career. And over that time my spend on AI has sky rocketed.
You don't have to believe me, it's obviously just anecdotal, but for anyone in the same position as me (And there really are lots of us), to claim model capability peaked in 2024 is just staggeringly dumb. It'd be like claiming electric cars peaked in 2008. I don't know how further to convey this to you.
It may very well be the case that the financial side is a bubble that horribly bursts. But the technology is real and the claims Ed has made are just wrong.
I'm the same as you sans 10 years of working experience, with similar experiences using LLMs. The distinction I would like to make is that it's not the models that are improving, but rather the infrastructure around them. It started with prompt engineering, then chain-of-thought, then mixture-of-experts, now harness engineering etc. These, I think are what's driving LLM complexity, not the model.
To bring back to your electric car analogy, the electric motor peaked early on, but was not quite useful as a car until advancements in battery density, charging technology, regenerative braking were made. These are all the infrastructure needed to improve the viability of owning electric cars, just like the infrastructure was needed to make the models useful.
To be honest, I'm much more excited to see development of infrastructure around local models than seeing the arms race between AI labs. At least the former benefits the end user in a transparent way.
1. Zitron claims model capability has peaked 2. Zitron claims AI lab growth (user and revenue) has stalled.
In the first case, TFA refutes by claiming 'wrong' repeatedly, which does not convince me of anything. If anything, Zitron is probably right in this regard, since the majority of progress in recent LLM tech has been setting up of guardrails to cajole the models using 'agents'.
In the second case, numbers are given showing growth, directly refuting Zitron. However, I'm giving Zitron the benefit of the doubt, given the old saying - market can remain irrational far longer than you can remain solvent. As long as people can be convinced that the sky is falling, rational predictions rarely pan out.