Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Of course I echo the universal: Opus 5 sucks. But what also sucks is still using 4.8. Are you all seeing this? It's like the older models got dumber just before Fable and Opus 5 were coming out? I've heard the theory that it's because 4.8 is getting put on older hardware? And I imagine that just ratchets down reasoning time then possibly?

Sometimes I'm just trying to sus out if I'm truly seeing things these days or going a little nuts :)



I have noticed very serious degradation in performance from Anthropic. I've switched away from Opus 5, but 4.8 is still much worse than how it was before Fable came out. It's constantly making what seems like obvious mistakes... I point them out, it's constantly apologizing.

I don't know why...

Is it A/B testing?

Is it load shedding?

Is it because I'm in Canada?

Is it because I'm not on the Claude Max plan?

Is it because I'm not paying via API?

Is it because I'm not paying via Bedrock?

Is it because the U.S. is worried people are distilling?

Is it because the U.S. wants to keep the top capability to themselves?

I think open models are the future. Anthropic is killing their reputation so fast. If they don't come clean I think they're cooked.


This is yet another reason why I think local models will win in the future. They're almost certainly A/B testing all sorts of opaque stuff that people have no clue about, hence the various 'How's Claude doing this session?' popups.

So what you are paying for may vary on a day by day basis, which is quite undesirable, even if their main goal is simply to make a better model. When it comes to a tool, I'd rather have consistent mediocrity than instability.


Yes. OpenAI is doing this as well, but they are much more transparent about it.


I definitely agree. I honestly cannot do tasks that require even minor complexity. Opus 5 keeps forgetting things in context as well and coding conventions. Really cannot build with CC without Fable.


Favourite conversation with Opus 5:

Me: why did you add HTMX?

Opus 5: you asked it twice.

Me: quote the exact sentence(s) where I asked it.

Opus 5: I can’t because you didn’t.


RIGHT!? The data suggests this isn't true but every fibre of my being is convinced that Opus 5 High was EXCELLENT at launch and has been lobotomised since then. My benchmark is Sol High. I've been using both consistently and either Sol High suddenly became MUCH more capable - and the data does not support that - or Opus 5 became much dumber. It's so bad I can't even use it anymore.


IMO Opus 5 wasn't that great at launch, but 4.8 was definitely downgraded just prior.

Agree that Sol is a great model. I find that it's improved a bit but I mostly attribute this to its eagerness to use the harness' memory features. (I'm using it in Hermes Agent, FWIW)


For me, Opus 5 mostly sucked because of its incomprehensible writing style. Having the system prompt focus on writing in terms that are easier to understand helped a bit


I'm not an expert by any means whatsoever but deploying models is not a straightforward task. There's a lot of levers to pull and I bet when models get "downgraded" to older hardware they do so WITHOUT the same stringent quality control of the output as they do when they release it.

I don't think it's something deliberately malicious like planned obsolescence but it's more like startup culture of "just make it fit in this sprint".


I attribute intent. These are for-profit companies with more leveraged capex than any other industry in history. They have enormous pressure to optimise their limited compute. Especially Anthropic, which is far more limited than OpenAI. Of course they're pulling those levers in the background to achieve "good enough." If they can reduce memory consumption by 20% and their metrics show a 4% loss of intelligence, they might very well pull that lever. These decisions compound.

Your explanation is probably the most likely and largest contributor. Anthropic states that Opsu 5 is a "pinned snapshot". They claim the weights and model configuration are not silently updated, BUT the surrounding serving infrastructure can change, including the request router, safety classifiers, and sampling logic. Anthropic has stated that if behaviour unexpectedly changes on a stable model ID, an infrastructure update is the most likely cause.

Further, "High" isn't a fixed amount of compute. Anthropic describes effort as a "behavioural signal" and not a token budget, with the model deciding how much thinking to do. So their "High" might be "Low" now, and we would never know.

Finally, I strongly suspect some quantisation or KV-cache compression is happening. Anthropic doesn't clearly delineate whether this would fall under the pinned weights and configuration, or the infrastructure, which almost certainly guarantees it's the latter. Forgetting earlier information, poor retrieval of details, contradicting previous conclusions, hallucination, degraded instruction-following, and losing the thread during complicated tasks are all symptoms of quantisation and compression.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: