When people talk about "real-world" in this context, the contrast is to artificial situations contrived as tests (e.g. benchmarks). This is blatantly clear to anyone actually engaging their brain.
Opus 5 does very well at the coding LLM benchmarks. In some cases even better than Fable 5.x. Yet in practice, in loads of real world situations where I've applied it against coding problems, it is vastly inferior to Fable.
When people talk about "real-world" in this context, the contrast is to artificial situations contrived as tests (e.g. benchmarks). This is blatantly clear to anyone actually engaging their brain.
Opus 5 does very well at the coding LLM benchmarks. In some cases even better than Fable 5.x. Yet in practice, in loads of real world situations where I've applied it against coding problems, it is vastly inferior to Fable.