You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
That's not the point, the point is that the company making the product is optimizing for the benchmark and/or the apparently idiosyncratic preferences of their own team, and not for the user experience of their paying customers.
Company can optimise for the benchmark (profit) while worsening the product. I think thebterm enshitification is used there. It appears that AI got it too
> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
There's a third perspective here: models are getting less useful, but overfitting to seeming useful to humans.
Imho, this is why analysis like TFA + third party cross-compatible harnesses (read: last mile UX) are so important to the leading labs optimizing for actual utility.
I'm suspicious enough of my subjective evaluation to believe a well-designed harness / verbiage could gaslight me into believing an objectively inferior model was superior. And at some point frontier labs are looking at the ROI of investing $1 in that vs actual model improvement.
I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged