I have a theory about this, what if we all became dumber after 4 months of heavy AI usage?
I remember how I enjoyed agents between December and February, something started changing around March.
I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6
You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
That's not the point, the point is that the company making the product is optimizing for the benchmark and/or the apparently idiosyncratic preferences of their own team, and not for the user experience of their paying customers.
Company can optimise for the benchmark (profit) while worsening the product. I think thebterm enshitification is used there. It appears that AI got it too
> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
There's a third perspective here: models are getting less useful, but overfitting to seeming useful to humans.
Imho, this is why analysis like TFA + third party cross-compatible harnesses (read: last mile UX) are so important to the leading labs optimizing for actual utility.
I'm suspicious enough of my subjective evaluation to believe a well-designed harness / verbiage could gaslight me into believing an objectively inferior model was superior. And at some point frontier labs are looking at the ROI of investing $1 in that vs actual model improvement.
The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?
A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.
I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.
I pay $200 a month at home for roughly the same amount of tokens I pay $3k for (or possibly more) at work. I doubt they are both profitable and sustainable.
Even that being the case, if the providers can squeeze more happy customers onto existing capacity they would likely act to increase profitability, no?
Benchmarks test whether models can pass exams with a right answer or a green test case. I don't think the models are getting dumber, but they're definitely getting more incomprehensible to talk to. I've noticed this happening almost as a step change with the overuse of words and tics, and so has the broader community apparently. We haven't all been getting dumb at the same rate.
>I thought models are getting dumber, but benchmarks were convincing opposite
>Opus 4.8 and Opus 5 seems worse models than Opus 4.6
After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.
And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.
In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.
There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.
I cannot speak to the benchmarks but we have many users, who have experience with each model, using it 95% of the week and we observe the same changes in the models as a whole.
In addition to that, while yes, 4.6 and 5.0 can solve problems, they do so differently. Sometimes 5.0 does better by a wide margin but that would be expected as they are supposed to be better.
It would describe the observed behavior. Especially if there were an internal quant/efficiency team that wasn't as diligent about regressions as the primary model team.
This is actually nonsense. More tokens means more capacity subscription and expense. Anthropic has enjoyed a high premium per million tokens because the quality per token was unusually high. Now it’s unusually low. This drives down the margin people will be willing to pay for the same number of tokens while driving up their capacity utilization. The economics are even worse for subscriptions.
Opus models have degraded rapidly since March, with each release being considerably less useful and considerably more verbose. The language is no so weirdly florid it’s difficult to understand, and its logical conclusions are almost always suspect. It goes off on clearly bizarre snipe hunts to the point it feels like I’m using a gpt 3 model at times. It’ll announce that it’s about to embark on building something then just return control to the user and wait. You can also tell perceptibly when they’re reducing model quality to load shed - it becomes stupider and stupider to the point you’re better off dumping state and switching to codex or just turning in for the day and hoping they secured more capacity tomorrow.
It’s an absolute race to the bottom with Anthropic on virtually every level. I’ve rarely seen a company so rapidly accumulate good will in the developer community as they did around 4.6 in December and January. By March, it was inconceivable to use anything else. 4.8 was a bit of a wake up call to not put all your harness eggs in one basket. 5 is straight up time to cancel territory.
I actually manually set my model back to the older versions to get anything serious done. More and more I use codex for anything non trivial.
This isn’t about avarice by the provide trying to get more tokens and more utilization. They’ve over subscribed for capacity as it is. If they can produce better quality for less tokens they can charge a higher margin and will be paid if, which is a better economic strategy overall. This is something else. I suspect it’s actually the opposite, they’re finding ways to cut capacity demand in ways that leads to worse behavior that leads to more capacity demands, worse output, worse quality, and worse margins, worse, worse, worse.
Just as I never saw a company accumulate such positive developer good will so fast, I’ve never seen one squander it so fast too.
Are you dumber? Can you do long division on paper?
You may say sure but why? We could cook over an open fire too but we have microwaves and stoves and restaurants and protein shakes.
Every generation since fire to bronze to internal combustion engines has adopted the new technology, integrated it so deeply into their lives that we recreate by going camping, disconnecting, or playing with toys that resemble the past era of forgotten tech.
I could stand Opus 4.6-4.8, I was impressed by the initial fable model. Codex 5.6 sol xhigh feels like the initial release of fable. Qwen 3.8 27b feels like using haiku or sonnet (I quickly stopped trying them).
I remember how I enjoyed agents between December and February, something started changing around March.
I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6