| |
The author argues that Opus 5, despite being more capable than earlier versions, feels worse to work with because it makes bold assumptions and reinterprets plans without asking for clarification, requiring constant oversight. They speculate this results from Anthropic's focus on benchmark performance and self-improving AI, which inherently rewards models that confidently handle ambiguity rather than seek clarification—traits that work well on benchmarks but are problematic for real-world coding tasks where human oversight and verification are essential.
Read Full Article →
← More Tech news