LLM Agents can Autonomously Hack Websites
2024/02/06 by Richard Fang, Rohan Bindu, Fang, Richard +7 · 8 voices · 48 citations
Business, Management and Accounting · Computer Science · #Blockchain Technology Applications and Security #Computer science #Computer security #Digital Rights Management and Security #FinTech, Crowdfunding, Digital Finance #Human–computer interaction #Internet privacy #World Wide Web #cs.AI #cs.CR
paper · pdf · doi:10.48550/arxiv.2402.06664
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In recent years, large language models (LLMs) have become increasingly capable and can now interact with tools (i.e., call functions), read documents, and recursively call themselves. As a result, these LLMs can now function autonomously as agents. With the rise in capabilities of these agents, recent work has speculated on how LLM agents would affect cybersecurity. However, not much is known about the offensive capabilities of LLM agents. In this work, we show that LLM agents can autonomously hack websites, performing tasks as complex as blind database schema extraction and SQL injections without human feedback. Importantly, the agent does not need to know the vulnerability beforehand. This capability is uniquely enabled by frontier models that are highly capable of tool use and leveraging extended context. Namely, we show that GPT-4 is capable of such hacks, but existing open-source models are not. Finally, we show that GPT-4 is capable of autonomously finding vulnerabilities in websites in the wild. Our findings raise questions about the widespread deployment of LLMs.
Cited by
- RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery
- CryptanalysisBench: Can LLMs do Cryptanalysis?
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
- Synthesizing Multi-Agent Harnesses for Vulnerability Discovery
- From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Quantifying Return on Security Controls in LLM Systems
- Fingerprint-Driven Automation: Coupling Reconnaissance with POC Verification
- Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
- When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
- Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
- OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Evaluating LLM Generated Detection Rules in Cybersecurity
- AI For Privacy in Smart Homes: Exploring How Leveraging AI-Powered Smart Devices Enhances Privacy Protection
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Agentic Discovery and Validation of Android App Vulnerabilities
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
- Prompt Injection 2.0: Hybrid AI Threats
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover
- TAI3: Testing Agent Integrity in Interpreting User Intent
- From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
- Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
- DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
- LLM Agents Should Employ Security Principles
- Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub
- Dynamic Risk Assessments for Offensive Cybersecurity Agents
- Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks
- Automated Profile Inference with Language Model Agents
- LLMs unlock new paths to monetizing exploits
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
- Concept-Level Explainability for Auditing & Steering LLM Responses
- RedTeamLLM: an Agentic AI framework for offensive security
- Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection
- From Texts to Shields: Convergence of Large Language Models and Cybersecurity
- Benchmarking Mythos-Linked Bug Rediscovery
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
- Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
- Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models
- Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
- "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
- ARCeR: an Agentic RAG for the Automated Definition of Cyber Ranges
Discussions
- LLM agents can autonomously hack websites [hn, 85 points, 21 comments]
- Die gute, hilfreiche KI ...... Kann autonom Websites hacken. 🤷♂️ arxiv.org/abs/2402.06664 [bsky, 3 points, 0 comments]
- LLM Agents can Autonomously Hack Websites [hn, 2 points, 0 comments]
- LLM Agents Can Autonomously Hack Websites (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- so... this just came out "LLM Agents can Autonomously Hack Websites" arxiv.org/abs/2402.06664 It's only going to get worse folks... [bsky, 1 points, 0 comments]
- Researchers trying offensive capabilities of LLM agents [bsky, 0 points, 0 comments]
- Well this looks like fun arxiv.org/abs/2402.06664 [bsky, 0 points, 0 comments]
- KI Sprachmodelle sind in der Lage, selbstständig Websites zu hacken, ohne Schwachstellen vorher zu kennen. "Wir zeigen, dass LLM-Agents selbstständig Websites hacken und so komplexe Aufgaben wie die [bsky, 0 points, 0 comments]
Related