| |
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Cognition released SWE-2, a 2.8-trillion-parameter mixture-of-experts model for software engineering that achieves 92.8 on Terminal-Bench 2.1 and 50.0 on FrontierCode 1.1, matching Claude Fable 5.1's performance at 64% lower cost. The model is built on Kimi K3 with Cognition's reinforcement learning post-training and is available through Devin Desktop and CLI, though weights remain proprietary. However, SWE-2 significantly underperforms on longer-horizon tasks (Terminal-Bench 4.0: 27.3 vs competitors' 55-58), indicating gaps remain in complex agentic work.
Read Full Article →
← More Tech news