| |
CursorBench 3.1
CursorBench 3.1 is a benchmark comparing various AI models on coding tasks, measuring both performance scores and cost-efficiency. Grok 4.6 leads in efficiency with a 70.8% score at just $2.81 per task, while Fable 5 and Opus 5 also rank highly, though at higher costs ranging from $8-$17 per task. The benchmark tracks multiple model versions across different configuration settings, with performance generally correlating to increased costs.
Read Full Article →
← More Tech news