Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I have not noticed this with Opus 4.6+. The result is usually not too far from what I would have written myself.


Opus 4.6 was the best model in the family, following two ones were seriously brain damaged to do well on benchmarks.


yeah those have been horrible


I wouldn't say terrible-terrible but only better at from-prompt-to-solution rather than interactive discussing and problem solving.

I tend to define it "better at solving, worse at assisting phenomenon". Which doesn't properly show on benchmarks that only focus on the solving part.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: