| |
Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
A browser-based game testing human oversight of AI coding agents revealed that players missed approximately 1 in 3 malicious commands, with an average accuracy of 66.3% across over 40,000 game runs. Credential exfiltration and scope violation threats were missed most frequently (33-35% miss rates), while obviously destructive commands like `rm -rf /` were caught more reliably (11.7% miss rate). The most-missed command was `npm run analyze`, with two-thirds of players approving it despite visible evidence in the history log that it contained malicious code, suggesting players often fail to carefully review agent action details before approval.
Read Full Article →
← More Tech news