I had the opposite experience. Claude models and its harness feel like they are set to eat tokens for everything, especially if it’s ultracode effort. I have seen it spawn 6 agents and eat my 4-hour quota right in front of me. Codex is on point and follows instructions well even with max effort. I like my analogy of Claude being garrulous and Codex being laconic.
If I see more limit hits in same sitting session, I have to rearrange my workflows. my experiments with qwen and deepseek have been good, cant wait to try glm and other models.
Just before I canceled my 20x Max sub I had a day where I had ~15% of my weekly I was trying to burn. I set it on ultracode and because I had set the max agents 32 for another project and forgot, the five research agents ended up spinning up a total of 26 subagents and burned through the remainder of my weekly in the span of 20 minutes before I noticed and shut it down.
Claude did the same thing to me the other day. I ran out of my 5 hr limit, and it went into my $100 credit they had given me. Multiple parallel agents. It ate it in no time flat.
When I checked, all that credit was gone, I still wasn't into the next 5 hours, and all the agents had failed, returning nothing.
I didn't even get anything for burning all that credit. If I had paid for it, I'd be very, very pissed.
Similar experience here. I had raised my cap to $20 because they had given me promo credit. The credit expired without me using it and I didn’t notice.
One fine day I was like, oh my weekly quota resets soon, I will kick off an expensive bug hunt. It launched parallel things and burned $10 in like a moment.
Here is an excerpt from the system prompt for UltraCode (Same for Fable,Opus,Sonnet):
"Ultracode. When a system-reminder confirms ultracode is on, that opt-in is standing: author and run a workflow for every substantive task by default. The goal is the most exhaustive, correct answer you can produce — token cost is not a constraint."