| |
Researchers evaluated five frontier AI models on their ability to fix 20 real-world security vulnerabilities (CVEs), finding that even the best-performing models only achieved a 50% success rate, with the most dangerous failure mode being patches that appear correct and pass tests while leaving vulnerabilities unfixed. The study also found that expensive models perform no better than cheaper alternatives within the same family, costing up to 12× more per run, and identified a critical risk: AI's false confidence in incomplete fixes could allow vulnerable code to ship undetected without human verification.
Read Full Article →
← More Tech news