| |
I built a vulnerable app and spent $1,500 seeing if LLMs could hack it
A security researcher spent $1,500 testing whether large language models could exploit a deliberately vulnerable book review app to find a hidden flag in private user data. GPT-5.5 was most successful with a 70% solve rate by identifying Firebase credentials in the app, while Deepseek V4 Pro achieved 30% success, and Claude, Gemini, and other models either failed entirely or were blocked by safety guardrails. The exploit demonstrates a real-world vulnerability class where hardened APIs paired with exposed Firebase credentials allow unauthorized database access.
Read Full Article →
← More Tech news