| |
Real-SWE is a new benchmark that evaluates frontier AI models on their ability to solve coding tasks from private, real-world enterprise codebases licensed from actual companies. The benchmark tests whether coding agents can handle the complexity of production environments, including proprietary systems, business-critical changes, and company-specific coding conventions. Initial results show Fable 5.1 leading with a 38.8% resolution rate, significantly outperforming other models tested.
Read Full Article →
← More Tech news