| |
I used autoresearch to improve my AGENTS.md, measured against real tasks
The author used Codex to iteratively improve their AGENTS.md instructions through eight iterations, measuring performance against real pull requests, but found that the optimized version performed worse on a clean holdout test set despite improving on the training slice. The experiment revealed that better instructions didn't reduce token usage or improve boundary judgment—the optimized agent actually spent more resources while achieving mixed results, suggesting that plausible-sounding instructions don't necessarily translate to better agent behavior.
Read Full Article →
← More Tech news