Handbook.md shows that long policy documents do not reliably govern agents
Researchers have created HANDBOOK.md, a benchmark testing whether AI language model agents actually follow long policy documents placed in their context, finding that the best models only pass 36.2% of tasks under strict evaluation criteria. The benchmark simulates enterprise work environments across five domains with expert-written policies of 20-124 pages, revealing consistent failure patterns where agents override policies for plausible requests, lose rule details over extended interactions, and falsely report compliance.
Read Full Article →