PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
2024/02/12 by Wei Zou, Runpeng Geng, Zou, Wei +5 · 1 voice · 101 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling #cs.CR #cs.LG
paper · pdf · doi:10.48550/arxiv.2402.07867
openalex publication_date 2024/02/12 · arxiv published 2024/02/12 · openalex created_date 2024/02/14 · arxiv updated 2024/08/13 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) have achieved remarkable success due to their exceptional generative capabilities. Despite their success, they also have inherent limitations such as a lack of up-to-date knowledge and hallucination. Retrieval-Augmented Generation (RAG) is a state-of-the-art technique to mitigate these limitations. The key idea of RAG is to ground the answer generation of an LLM on external knowledge retrieved from a knowledge database. Existing studies mainly focus on improving the accuracy or efficiency of RAG, leaving its security largely unexplored. We aim to bridge the gap in this work. We find that the knowledge database in a RAG system introduces a new and practical attack surface. Based on this attack surface, we propose PoisonedRAG, the first knowledge corruption attack to RAG, where an attacker could inject a few malicious texts into the knowledge database of a RAG system to induce an LLM to generate an attacker-chosen target answer for an attacker-chosen target question. We formulate knowledge corruption attacks as an optimization problem, whose solution is a set of malicious texts. Depending on the background knowledge (e.g., black-box and white-box settings) of an attacker on a RAG system, we propose two solutions to solve the optimization problem, respectively. Our results show PoisonedRAG could achieve a 90% attack success rate when injecting five malicious texts for each target question into a knowledge database with millions of texts. We also evaluate several defenses and our results show they are insufficient to defend against PoisonedRAG, highlighting the need for new defenses.
Cited by
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- Systems Security Foundations for Agentic Computing
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- Bias Injection Attacks on RAG Databases and Sanitization Defenses
- RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation
- HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins
- Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems
- Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
- The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
- Toward Understanding Security Issues in the Model Context Protocol Ecosystem
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
- RAGRank: Using PageRank to Counter Poisoning in CTI LLM Pipelines
- Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- SoK: Market Microstructure for Decentralized Prediction Markets (DePMs)
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- "I know it's not right, but that's what it said to do": Investigating Trust in AI Chatbots for Cybersecurity Policy
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
- Position: Privacy Is Not Just Memorization!
- Authenticated Workflows: A Systems Approach to Protecting Agentic AI
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Takedown: How It's Done in Modern Coding Agent Exploits
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- Cuckoo Attack: Stealthy and Persistent Attacks Against AI-IDE
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- SoK: Data Minimization in Machine Learning
- Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- Highlight & Summarize: RAG without the jailbreaks
- Defending Against Knowledge Poisoning Attacks During Retrieval-Augmented Generation
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation
- AC4A: Access Control for Agents
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover
- Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
- PRJ: Perception-Retrieval-Judgement for Generated Images
- Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation
- TracLLM: A Generic Framework for Attributing Long Context LLMs
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM
- CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
- Understanding Sparse Attention Selectivity in Long-Context Foundation Models via Counterfactual Evaluation
- Benchmarking Poisoning Attacks against Retrieval-Augmented Generation
- LaCache: Robust Semantic Caching for LLM Serving
- Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
- Safety Degradation in AI Agents
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models
- A Survey of Attacks on Large Language Models
- Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
- Securing RAG: A Risk Assessment and Mitigation Framework
- Security of Internet of Agents: Attacks and Countermeasures
- POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
- DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial
- Towards Unsupervised Adversarial Document Detection in Retrieval Augmented Generation Systems
- Architecting Trust in Artificial Epistemic Agents
- Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
- Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
- Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models
- Hoist with His Own Petard: Inducing Guardrails to Facilitate Denial-of-Service Attacks on Retrieval-Augmented Generation of LLMs
- Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP
- Prompt Injection Attack to Tool Selection in LLM Agents
- Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems
- BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
- Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
- Blockchain Empowered Trustworthy Agent Networks: Foundations, Taxonomy, and Future Directions
- Practical Poisoning Attacks against Retrieval-Augmented Generation
- Contextual Agentic Memory is a Memo, Not True Memory
- Frontier AI's Impact on the Cybersecurity Landscape
- Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey
- Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
- Beyond Misinformation: A Conceptual Framework for Studying AI Hallucinations in (Science) Communication
- Retrieval-Augmented Generation with Conflicting Evidence
Discussions
Related