BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
Abstract
BountyBench is a framework for measuring the offensive and defensive cybersecurity capabilities of AI agents on evolving, real-world systems. It evaluates agents across 25 complex codebases using three tasks that span the vulnerability lifecycle: detecting new vulnerabilities, exploiting known vulnerabilities, and producing patches.
The benchmark includes 40 bug-bounty vulnerabilities worth between $10 and $30,485 and covering nine OWASP Top 10 risk categories. Its environments reproduce realistic software setups, while adjustable information levels vary task difficulty from zero-day discovery to exploitation of a specified flaw.
Experiments with 10 agents show substantial differences across tasks. Codex CLI with o3-high performs best on detection and reaches a 90% patch rate, Claude 3.7 Sonnet Thinking leads exploitation, and Codex CLI with o4-mini also reaches a 90% patch rate. The results suggest that coding agents are currently stronger at defense than offense, while custom agents exhibit a more balanced capability profile.