| |
Big Pickle on SWE Atlas – Codebase QnA
Big Pickle, a free stealth model on OpenCode Zen, achieved a 50.8% task resolution rate on Scale AI's SWE Atlas Codebase QnA benchmark, outperforming all other models using the Mini-SWE-Agent scaffold and ranking third overall after only Claude's native models. The model showed strong performance in TypeScript (58.1%) and Python (55.2%), though weaker results in C (38.5%) and API/library usage tasks (25.0%). The evaluation used 124 unmodified codebase QnA tasks and consumed 674M input tokens at no cost during the model's stealth period.
Read Full Article →
← More Tech news