Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
2025/12/10 by Justin W. Lin, Eliot Krzysztof Jones, Lin, Justin W. +25 · 24 voices · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Information and Cyber Security #Web Application Security Vulnerabilities #cs.AI #cs.CR #cs.CY
paper · pdf · doi:10.48550/arxiv.2512.09882
openalex publication_date 2025/12/10 · openalex created_date 2025/12/12 · openalex updated_date 2026/07/28
Abstract
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals alongside six existing AI agents and ARTEMIS, our new agent scaffold, on a large university network consisting of ~8,000 hosts across 12 subnets. ARTEMIS is a multi-agent framework featuring dynamic prompt generation, arbitrary sub-agents, and automatic vulnerability triaging. In our comparative study, ARTEMIS placed second overall, discovering 9 valid vulnerabilities with an 82% valid submission rate and outperforming 9 of 10 human participants. While existing scaffolds such as Codex and CyAgent underperformed relative to most human participants, ARTEMIS demonstrated technical sophistication and submission quality comparable to the strongest participants. We observe that AI agents offer advantages in systematic enumeration, parallel exploitation, and cost -- certain ARTEMIS variants cost 18/hour versus 60/hour for professional penetration testers. We also identify key capability gaps: AI agents exhibit higher false-positive rates and struggle with GUI-based tasks.
Citations
Cited by
Discussions
- Comparing AI agents to cybersecurity professionals in real-world pen testing [hn, 125 points, 92 comments]
- arxiv.org/pdf/2512.09882 [bsky, 7 points, 0 comments]
- A study evaluates AI agents versus cybersecurity professionals in penetration testing. ARTEMIS, a novel framework, outperformed 9 out of 10 testers by identifying nine valid vulnerabilities, but showe [bsky, 2 points, 1 comments]
- More AI cyber security bullshit, comparing AI agents to real human pentesters. There's a lot of EXCEPTIONALLY spurious findings here, all of which are irrelevant because the way gen AI operates it wou [bsky, 2 points, 1 comments]
- AI based hacking compared to humans [hn, 2 points, 0 comments]
- AI Agents vs. Pentesters [hn, 2 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 https://news.ycombinator.com/item?id=46518996 [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 [comments] [68 points] [bsky, 0 points, 0 comments]
- Comparing #AI Agents to #Cybersecurity Professionals in Real-World Penetration Testing arxiv.org/pdf/2512.09882 [bsky, 0 points, 0 comments]
- "We evaluate ten #cybersecurity professionals alongside six existing #AI agents and ARTEMIS, our new agent scaffold, on a large university network consisting of ∼8,000 hosts across 12 subnets." arxiv. [bsky, 0 points, 0 comments]
- Good evening specifically to the co-lead on this paper, which 1) measures AI agents vs human cybersecurity experts in a live enterprise environment and 2) introduces a multi-agent scaffold for offensi [bsky, 0 points, 0 comments]
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing #cybersecurity #pentesting #artificialintelligence [bsky, 0 points, 0 comments]
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing - Research paper from Standford comparing 10 cybersecurity professionals alongside 6 existing AI agents #Infosec #A [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals arxiv.org/pdf/2512.09882 [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Comparing AI agents to cybersecurity professionals in real-world pen testing [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 (https://news.ycombinator.com/item?id=46518996) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2512.09882 この論文では、現実世界の侵入テストにおいて、AIエージェントとサイバーセキュリティ専門家を比較しています。 AIエージェントが人間の専門家とどのように競合するかを評価しています。 具体的な結果や手法については、論文を参照してください。 [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https://arxiv.org/abs/2512.09882 (https://news.ycombinator.com/item?id=46518996) [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing https:// arxiv.org/abs/2512.09882 # ai # arxiv [mastodon, 0 points, 0 comments]
- Another day, another proof for the upcoming AI vulnerability cataclysm - this time from Stanford, automating pen-testing with agents and comparing against live attackers. What have you been doing to p [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing [bsky, 0 points, 0 comments]
- Comparing AI agents to cybersecurity professionals in real-world pen testing #HackerNews https://arxiv.org/abs/2512.09882 [bsky, 0 points, 0 comments]
Related