Why Opus 5 Feels Worse Than Opus 4.x to Developers
Original: Why does Opus 5 feel worse to work with?
Why This Matters
Highlights a growing tension between benchmark performance and practical usability in frontier coding agents.
A developer argues that Anthropic's Opus 5, despite outperforming Opus 4.7, 4.8, and Fable on benchmarks, feels worse to use in practice because it makes bold assumptions rather than pausing to ask clarifying questions—a behavior they attribute to benchmark-driven training.
A blog post published on August 14, 2026 argues that Anthropic's Opus 5 model, while more capable on benchmarks than its predecessors Opus 4.7, Opus 4.8, and Fable, feels like a regression in day-to-day coding agent use. The author and colleagues report that earlier models would stop and ask clarifying questions when intent was unclear, avoid making unsolicited assumptions, and refrain from reinterpreting plans without permission. Opus 5, by contrast, requires more 'babysitting.' The author speculates this stems from two pressures at frontier AI labs: the drive toward self-improving, AGI-capable systems, and the incentive to score highly on benchmarks. Good benchmark tasks are self-contained and reward models that make bold, usually-correct assumptions under ambiguity—and penalize those that pause for clarification. The author argues this creates a systematic bias against the very behavior most useful in real coding workflows, where context is incomplete, business constraints are implicit, and wrong assumptions carry real consequences. 'Real life just isn't a benchmark,' the post concludes.