| |
Can open-source prompt-injection detectors catch realistic AI agent attacks?
A benchmark testing 10 open-source prompt-injection detectors against 629 realistic AI agent attacks found that none effectively catch most attacks without also blocking legitimate traffic. The best detector caught 51% of attacks while causing only 2% false positives, while Meta's Prompt Guard 2 caught just 1%, and detectors fail in three ways: not recognizing attack wording, losing detection in surrounding context, or flagging nearly all traffic as malicious.
Read Full Article →
← More Tech news