Benchmarking Opus 5 on SlopCodeBench - AllTheNews.today
Benchmarking Opus 5 on SlopCodeBench

Benchmarking Opus 5 on SlopCodeBench

Researchers benchmarked Claude's Opus 5 model on SlopCodeBench, a new long-horizon coding benchmark that tests a model's ability to maintain code quality as requirements evolve over time rather than revealing the entire problem upfront. Opus 5 achieved a 24% pass rate on the tested subset, only marginally higher than the original paper's reported 17% for Opus 4.6, while all tested models showed significant increases in code complexity and verbosity throughout the challenges. The results suggest that current AI models cannot reliably handle real-world software engineering tasks independently without human guidance or steering.
Read Full Article →
github.com
← Back to Latest