Formalizing and Benchmarking Prompt Injection Attacks and Defenses
2023/10/19 by Yupei Liu, Yuqi Jia, Liu, Yupei +7 · 94 citations
Computer Science · #Security and Verification in Computing #Advanced Malware Detection Techniques #Internet Traffic Analysis and Secure E-voting
paper · pdf · doi:10.48550/arxiv.2310.12815
Abstract
A prompt injection attack aims to inject malicious instruction/data into the input of an LLM-Integrated Application such that it produces results as an attacker desires. Existing works are limited to case studies. As a result, the literature lacks a systematic understanding of prompt injection attacks and their defenses. We aim to bridge the gap in this work. In particular, we propose a framework to formalize prompt injection attacks. Existing attacks are special cases in our framework. Moreover, based on our framework, we design a new attack by combining existing ones. Using our framework, we conduct a systematic evaluation on 5 prompt injection attacks and 10 defenses with 10 LLMs and 7 tasks. Our work provides a common benchmark for quantitatively evaluating future prompt injection attacks and defenses. To facilitate research on this topic, we make our platform public at https://github.com/liu00222/Open-Prompt-Injection.
Cited by
- Agent Data Injection Attacks are Realistic Threats to AI Agents
- Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- PINA: Prompt Injection Attack Against Navigation Agents
- ContextLeak: Auditing Leakage in Private In-Context Learning Methods
- Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- Phishing Email Detection Using Large Language Models
- Chasing Shadows: Pitfalls in LLM Security Research
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
- Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection
- "MCP Does Not Stand for Misuse Cryptography Protocol": Uncovering Cryptographic Misuse in Model Context Protocol at Scale
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Many-to-One Adversarial Consensus: Exposing Multi-Agent Collusion Risks in AI-Based Healthcare
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- Bias Injection Attacks on RAG Databases and Sanitization Defenses
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Building Browser Agents: Architecture, Security, and Practical Solutions
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
- RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
- Toward Understanding Security Issues in the Model Context Protocol Ecosystem
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Secure Retrieval-Augmented Generation against Poisoning Attacks
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Soft Instruction De-escalation Defense
- Prompt injections as a tool for preserving identity in GAI image descriptions
- Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- PromptLocate: Localizing Prompt Injection Attacks
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
- Text Prompt Injection of Vision Language Models
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Exploiting Web Search Tools of AI Agents for Data Exfiltration
- Authenticated Workflows: A Systems Approach to Protecting Agentic AI
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
- Imperceptible Jailbreaking against Large Language Models
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
- Fingerprinting LLMs via Prompt Injection
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System Vectors
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
- Investigating Security Implications of Automatically Generated Code on the Software Supply Chain
- Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
- DecipherGuard: Understanding and Deciphering Jailbreak Prompts for a Safer Deployment of Intelligent Software Systems
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- Free-MAD: Consensus-Free Multi-Agent Debate
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
- SoK: Large Language Model Copyright Auditing via Fingerprinting
- Retrieval-Augmented Review Generation for Poisoning Recommender Systems
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- Interface on demand: Towards AI native Control interfaces for 6G
- When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- Provably Secure Retrieval-Augmented Generation
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era
- Understanding the Supply Chain and Risks of Large Language Model Applications
- PAuth - Precise Task-Scoped Authorization For Agents
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
- When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- Defending Against Prompt Injection With a Few DefensiveTokens
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover
- How Not to Detect Prompt Injections with an LLM
- Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
- DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Linearly Decoding Refused Knowledge in Aligned Language Models
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Advanced Applications of Generative AI in Actuarial Science: Case Studies Beyond ChatGPT
Related