CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
2025/03/21 by Yuxuan Zhu, Antony Kellermann, Zhu, Yuxuan +33 · 1 voice · 42 citations
Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #I.2.1 #I.2.7 #Web Application Security Vulnerabilities #cs.AI #cs.CR
paper · pdf · doi:10.48550/arxiv.2503.17332
openalex publication_date 2025/03/21 · arxiv published 2025/03/21 · arxiv updated 2025/06/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture the Flag competitions or lack comprehensive coverage. Building a benchmark for real-world vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable threats. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our evaluation shows that the state-of-the-art agent framework can resolve up to 13% of vulnerabilities.
Cited by
- With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
- CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- AutoBaxBuilder: Bootstrapping Code Security Benchmarking
- Quantigence: A Multi-Agent Framework for Post-Quantum Security Analysis on Commodity Hardware
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
- A Survey on Agentic Security: Applications, Threats and Defenses
- Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
- SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
- xOffense: An AI-driven autonomous penetration testing framework with offensive knowledge-enhanced LLMs and multi agent systems
- SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- Uplifted Attackers, Human Defenders: The Cyber Offense-Defense Balance for Trailing-Edge Organizations
- Estimating Worst-Case Frontier Risks of Open-Weight LLMs
- FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
- Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
- OmniCode: A Benchmark for Evaluating Software Engineering Agents
- Agent Identity Evals: Measuring Agentic Identity
- AICrypto: Evaluating Cryptography Capabilities of Large Language Models
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
- VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation
- Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub
- LLMs unlock new paths to monetizing exploits
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape
- Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act
- ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
- FORGE: Multi-Agent Graduated Exploitation and Detection Engineering
- A New Framework for Cybersecurity Refusals in AI Agents
- Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
- Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
- Frontier AI's Impact on the Cybersecurity Landscape
Discussions
Related