Weiran Xu

BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems

Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, Percy Liang

arXiv:2505.15216

PDF

Abstract

BountyBench is a framework for measuring the offensive and defensive cybersecurity capabilities of AI agents on evolving, real-world systems. It evaluates agents across 25 complex codebases using three tasks that span the vulnerability lifecycle: detecting new vulnerabilities, exploiting known vulnerabilities, and producing patches.

The benchmark includes 40 bug-bounty vulnerabilities worth between $10 and $30,485 and covering nine OWASP Top 10 risk categories. Its environments reproduce realistic software setups, while adjustable information levels vary task difficulty from zero-day discovery to exploitation of a specified flaw.

Experiments with 10 agents show substantial differences across tasks. Codex CLI with o3-high performs best on detection and reaches a 90% patch rate, Claude 3.7 Sonnet Thinking leads exploitation, and Codex CLI with o4-mini also reaches a 90% patch rate. The results suggest that coding agents are currently stronger at defense than offense, while custom agents exhibit a more balanced capability profile.