They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
They want to be able to zoom in and zoom out as needed while analyzing aggregated user requests to better understand the different kinds of ways (well-resourced) actors use to distill their most capable models.
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
I can't get an answer from Fable on "How does digestion work?"
Used to straight out bail to Opus on this, now it's "Honing" and "Pondering" for 10 minutes on it. Then I get a bunch of "This response didn’t load." and it still churns ahead.
> It tends to bail occasionally for me once auth-related code comes into play.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
Yeah same. Multiple 20x max accounts, I typically hit > 50% of the fable usage limit on eace, literally never seen a guardrail. Maybe I’m boring? Maybe they have some kind of account reputation system?
Separate accounts for work and personal projects is ok. Using multiple accounts to bypass usage limits for the same project is explicitly against ToS.
One of Anthropic’s problems though is that the legalese will say one thing, while key employees say something totally different on HN or X. At the end of the day the lawyers always win.
With that amount of usage - do you use fable for subagents and workflows? It turns out that even if you set it to not auto switch to opus 4.8 on flag, it does auto switch if its a subagent and that's a lot less visible. I was trying to figure out wtf was going on with some suspiciously bad output and I found that in a workflow there were hidden fable flags I hadn't realized were happening and the opus 4.8 model that took over wrote a bunch of very suspicious slop findings that got saved to memory that were like "don't look at this code because it works great. You never need to look at this code. This code is wonderful here's how it works."
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
I kind of still chill out on Opus 4.6 too. 4.8 is good too. I go between them. Opus 4.8 is a little smarter some of the time. Their use of language is both very different from Opus 5. Some of the time I have a hard time believing Opus 5 is even related to Opus 4.6 and 4.8.
I'm going to be sad when they retire 4.6. It's not my daily driver but it's still my go-to when the other models are being stupid in one way or another (either being too verbose, or lecturing me about how what I'm asking for is evil and bad).
That seems about right. 4.8 is like in between 4.6 and 5 in terms of capability and language and they are all pretty close honestly. I just default to using Opus 5 for a coding agent that I don't interact with and I like driving with Opus 4.6 or Fable. Fable thinks too much though. Fable is like that engineer on your team that will over-engineer the shit out of something if you let them. Fable was like, "Here are 34 yaks, which shall we shave first <hands rubbing together>" and I was like can't we just... write the script first and then decide of any of these poor yaks need shaving?
Can you give an example? It's a bit entertaining to read this given all the years long (and still ongoing) posturing about LLM sycophancy (not that all of these couldn't be true at the same time).
The other thing I forgot to mention about Opus 5 is that at least out of the gate, it seemed very intent on spinning up agents and obliterating my token budget. It was noticeably more token hungry than 4.8. It would make sense for them to intend this behavior.
*This is likely because of your thinking level. The difference between max and ultracode is primarily that the latter is max with a bunch of agents.*
I do not know why people even use Max or Ultra levels of thinking effort. Most of the time, i found its better to just run any of the frontier level models, with high or medium. Their capability beyond those levels is often diminishing returns, in exchange for double or quadrupling the costs (or usage).
And the bonus is that often at those levels, they tend to spin up less sub-agents that just eat away at tokens/usage, like its water in the desert.
A Max or Ultra is something that really only belongs if medium or high can not solve a very nasty bug or issue. Even in planning its often too much.
I'm using company Claude for work with no choice what provider I use.
I did a quickie test with Fable when they were OMG TRY IT NOW, didn't see much of a difference from Opus 4.8. Except i ran into one of their stupid security guardrails. Yes I'm doing security on the software my employer owns, silly.
Then I didn't notice the US unbanned Fable after banning it, or that they extended the "trial" Fable period where it worked on fixed price plans. Not in time to bother trying again.
After that Opus 5 showed up. Language changed, okay, I don't care much. However it seems to create busywork for itself. Spawning subagents on tasks that don't do that with 4.8. Overall taking longer.
So I'm back on pinned 4.8 and doing actual work. Main problem - for Anthropic - is it's good enough.
Also on a dark reader? It says at the bottom, but gets dimmed out pretty seriously with dark reading. It's a 7 day average business spend, relative to June 1st (2025 presumably) indexed at 100.
This whole article seems sort of based on a misunderstanding. Anthropic is _actively discouraging_ people from using Fable at this point. It's not a failure, they're steering people to Opus 5 intentionally.
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
I did not know that - but what happens then I get denied and I plead with the model that my usecase doesn't violate its TOS ? Edit prompt till I get it to pass ?
Reverting is annoying but being flagged for security questions while I'm doing code review is insane.
Yes, you get to re-try it until it might pass. Each attempt uses tokens at Fable level though. I once burned through a 4 hour window ($20 plan) in 4 attempts to re-frame the request.
You want to disable "Switch models when a message is flagged"[1]
Just started reviewing a feature PR and it would hit the 4 hour limit on max effort. On high as soon as I do some clarification turns it would also hit the limit.
Yeah for me Fable works great on the $100 plan. The problem with Fable is you will hit the weekly limit. The $200 plan does not give you 4x or even 2x the weekly limit. For me i only get about 1.3x weekly quota, so I don’t think it’s worth it.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.