| |
DoGBench: The first user-facing docs generation benchmark. No model scores >50%
DoGBench is the first benchmark evaluating AI agents' ability to write and maintain user-facing software documentation, with 292 tasks from real open-source projects. No model scored above 50% on the evaluation, with the highest composite score of 54.8% held by a cloud agent, revealing that agents struggle both with deciding when documentation needs updating and with producing documentation that enables task completion despite polished writing. Common failures include missing prerequisites or steps (45.5%), technical inaccuracies (36.6%), and omitted key information (32.5%).
Read Full Article →
← More Tech news