| |
Terminal-Bench-Science is a new benchmark developed by Stanford researchers to evaluate AI agents on real scientific research workflows across multiple disciplines, with 70 tasks in its first release. The strongest model tested, Claude Opus 5, achieved only a 30% resolution rate, demonstrating significant gaps in AI's ability to assist with scientific work. The benchmark aims to drive development of AI agents that can handle technically demanding research tasks, allowing scientists to focus on higher-level aspects like hypothesis formation and result interpretation.
Read Full Article →
← More Tech news