Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
2023/02/23 by Kai Greshake, Greshake, Kai, Sahar Abdelnabi +9 · 10 voices · 445 citations
Computer Science · Social Sciences · #Access Control and Trust #Adversarial system #Artificial intelligence #Computer science #Computer security #Exploit #Interface (matter) #Key (lock) #Software deployment #Software engineering #Topic Modeling #Web Application Security Vulnerabilities #cs.AI #cs.CL #cs.CR #cs.CY
paper · pdf · doi:10.48550/arxiv.2302.12173
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/02/23 · openalex created_date 2023/02/25 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) are increasingly being integrated into applications, with versatile functionalities that can be easily modulated via natural language prompts. So far, it was assumed that the user is directly prompting the LLM. But, what if it is not the user prompting? We show that LLM-Integrated Applications blur the line between data and instructions and reveal several new attack vectors, using Indirect Prompt Injection, that enable adversaries to remotely (i.e., without a direct interface) exploit LLM-integrated applications by strategically injecting prompts into data likely to be retrieved at inference time. We derive a comprehensive taxonomy from a computer security perspective to broadly investigate impacts and vulnerabilities, including data theft, worming, information ecosystem contamination, and other novel security risks. We then demonstrate the practical viability of our attacks against both real-world systems, such as Bing Chat and code-completion engines, and GPT-4 synthetic applications. We show how processing retrieved prompts can act as arbitrary code execution, manipulate the application's functionality, and control how and if other APIs are called. Despite the increasing reliance on LLMs, effective mitigations of these emerging threats are lacking. By raising awareness of these vulnerabilities, we aim to promote the safe and responsible deployment of these powerful models and the development of robust defenses that protect users from potential attacks.
Cited by
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- Protocol-Level Attacks on Agentic Commerce Platforms: A Cross-Platform Taxonomy, AIP-Bench, and Unified Defense
- ToolGuardian: Declarative Security for AI Agent-Tool Interactions
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
- Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
- Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
- Twin Agent: Context Residual Compression for Privilege Separated Agents
- Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
- Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
- Data Leakage Prevention in Agentic Applications via Preemptive Hardening
- PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents
- Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
- Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
- RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
- Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
- Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
- Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings
- DisarmRAG: Stealthy Retriever-Centric Poisoning to Disable Self-Correction in Retrieval-Augmented Generation (Extended Version)
- How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice
- AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
- Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation Evaluation
- Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents
- Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
- It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance
- Ghost Vectors: Soft-Deleted Embeddings Remain Reconstructible in HNSW Vector Databases
- Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
- The End of Code Review: Coding Agents Supersede Human Inspection
- NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
- Incomplete Prompt Jailbreaks in Large Language Models
- Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- TopoGuard: Graph Theory Based Defenses Against Split-Knowledge Attacks on RAG
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
- Distributional AGI Safety
- A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
- Securing AI Agents with Information-Flow Control
- Exploiting large language models in peer review: indirect prompt injection attacks and integrity probes
- Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
- Threat Modeling for AI: The Case for an Asset-Centric Approach
- Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
- Prismata: Confining Cross-Site Prompt Injection in Web Agents
- Agent Security is a Systems Problem
- Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing
- Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
- Agent Data Injection Attacks are Realistic Threats to AI Agents
- Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric
- Where Is the Cost of Third-Party API Routers in Agentic Software Development?
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
- TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
- Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- The Missing Layer: Specification Infrastructure for AI Oversight
- False Prophets: On the Security of World Models in Agentic Systems
- How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
- Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents
- PINA: Prompt Injection Attack Against Navigation Agents
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- DREAM: Dynamic Red-teaming across Environments for AI Models
- SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
- IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol
- Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
- Detecting Prompt Injection Attacks Against Application Using Classifiers
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
- How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation
- Phishing Email Detection Using Large Language Models
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data
- Black-Box Behavioral Distillation Breaks Safety Alignment in Medical LLMs
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
- MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection
- Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- Personalizing Agent Privacy Decisions via Logical Entailment
- Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
- Systems Security Foundations for Agentic Computing
- Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
- Bias Injection Attacks on RAG Databases and Sanitization Defenses
- Toward a Safe Internet of Agents
- An Empirical Study on the Security Vulnerabilities of GPTs
- AgentShield: Make MAS more secure and efficient
- BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
- AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
- Strategic Decision Framework for Enterprise LLM Adoption
- MURMUR: Using cross-user chatter to break collaborative language agents in groups
- Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
- Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments
- SnapAudit: Active Auditing of Differentially Private In-Context Learning via Snapshot-Based Simulation
- Securing Generative AI in Healthcare: A Zero-Trust Architecture Powered by Confidential Computing on Google Cloud
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins
- ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- Caption Injection for Optimization in Generative Search Engine
- Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents
- Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
- Distributing Security Controls Through Harness Engineering
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- GPT-Red: Automated Red Teaming via Self-Play at Scale
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- Safety from Honesty in a Disinterested AI Predictor
- The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
- The Consensus Trap: Rescuing Multi-Agent LLMs from Adversarial Majorities via Token-Level Collaboration
- Toward Understanding Security Issues in the Model Context Protocol Ecosystem
- The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
- MCPGuard : Automatically Detecting Vulnerabilities in MCP Servers
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
- AgentBound: Securing Execution Boundaries of AI Agents
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge
- Soft Instruction De-escalation Defense
- AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices
- Defending Against Prompt Injection with DataFilter
- The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
- Black-box Optimization of LLM Outputs by Asking for Directions
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
- Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- PromptLocate: Localizing Prompt Injection Attacks
- Attacks by Content: Automated Fact-checking is an AI Security Issue
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- A Vision for Access Control in LLM-based Agent Systems
- BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- Red-Teaming the Agentic Red-Team
- Who Owns This Agent? Tracing AI Agents Back to Their Owners
- SoK: Security of Autonomous LLM Agents in Agentic Commerce
- Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Exploiting Web Search Tools of AI Agents for Data Exfiltration
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Chain-of-Trigger: An Agentic Backdoor that Paradoxically Enhances Agentic Robustness
- Intelligent AI Delegation
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- A Survey on Agentic Security: Applications, Threats and Defenses
- Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs
- Adversarial Reinforcement Learning for Large Language Model Agent Safety
- Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours
- Better Privilege Separation for Agents by Restricting Data Types
- Fingerprinting LLMs via Prompt Injection
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- Takedown: How It's Done in Modern Coding Agent Exploits
- Incentive-Aligned Multi-Source LLM Summaries
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
- FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
- Investigating Security Implications of Automatically Generated Code on the Software Supply Chain
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
- What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
- AI Agents May Always Fall for Prompt Injections
- Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
- Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
- PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
- Enterprise AI Must Enforce Participant-Aware Access Control
- AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- MillStone: How Open-Minded Are LLMs?
- Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
- From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
- Free-MAD: Consensus-Free Multi-Agent Debate
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
- SoK: Security and Privacy of AI Agents for Blockchain
- Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
- EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints
- Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment
- Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
- A Survey: Towards Privacy and Security in Mobile Large Language Models
- The Aegis Protocol: A Foundational Security Framework for Autonomous AI Agents
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
- Reliable Weak-to-Strong Monitoring of LLM Agents
- UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- Effective Red-Teaming of Policy-Adherent Agents
- Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous
- Role-Augmented Intent-Driven Generative Search Engine Optimization
- Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment
- AI Security Map: Holistic Organization of AI Security Technologies and Impacts on Stakeholders
- When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
- Evo-MARL: Co-Evolutionary Multi-Agent Reinforcement Learning for Internalized Safety
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- The SMeL Test: A simple benchmark for media literacy in language models
- A Survey on Data Security in Large Language Models
- LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks
- Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
- Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
- Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding
- Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models
- Does More Inference-Time Compute Really Help Robustness?
- Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
- Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Definition as a Code, and Zero-Trust Access
- PAuth - Precise Task-Scoped Authorization For Agents
- The Controllability Trap: A Governance Framework for Military AI Agents
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild
- WebGuard: Building a Generalizable Guardrail for Web Agents
- When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs
- Benchmarking LLM Privacy Recognition for Social Robot Decision Making
- TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
- Adversarial machine learning :
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- Design Patterns for Securing LLM Agents against Prompt Injections
- Defending Against Prompt Injection With a Few DefensiveTokens
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
- The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover
- How Not to Detect Prompt Injections with an LLM
- Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
- Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
- Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion
- CyberNeuro: A Privacy-Preserving Agentic Workbench for Cohort-Scale Neuroimage and Clinical Data Analysis
- The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- Control at Stake: Evaluating the Security Landscape of LLM-Driven Email Agents
- Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
- GPT, But Backwards: Exactly Inverting Language Model Outputs
- Enhancing LLM Agent Safety via Causal Influence Prompting
- Linearly Decoding Refused Knowledge in Aligned Language Models
- GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- Securing AI Systems: A Guide to Known Attacks and Impacts
- More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Enhancing Security in LLM Applications: A Performance Evaluation of Early Detection Systems
- AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
- Context manipulation attacks : Web agents are susceptible to corrupted memory
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents
- Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks
- Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
- Defending against Indirect Prompt Injection by Instruction Detection
- Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
- Control Tax: The Price of Keeping AI in Check
- Normative Conflicts and Shallow AI Alignment
- Detection Method for Prompt Injection by Integrating Pre-trained Model and Heuristic Feature Engineering
- Sentinel: SOTA model to protect against prompt injections
- Maris: A Formally Verifiable Privacy Policy Enforcement Paradigm for Multi-Agent Collaboration Systems
- Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
- Misalignment or misuse? The AGI alignment tradeoff
- TracLLM: A Generic Framework for Attributing Long Context LLMs
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- Distinguishing Autonomous AI Agents from Collaborative Agentic Systems: A Comprehensive Framework for Understanding Modern Intelligent Architectures
- ACCESS DENIED INC: The First Benchmark Environment for Sensitivity Awareness
- Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
- Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
- Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
- Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems
- When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
- Automatic Programming: Large Language Models and Beyond
- RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM
- Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
- Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
- When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
- The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
- When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
- Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
- Security Concerns for Large Language Models: A Survey
- Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- Tool Preferences in Agentic LLMs are Unreliable
- CAIN: Hijacking LLM-Humans Conversations via Malicious System Prompts
- Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
- In-Context Watermarks for Large Language Models
- Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
- EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection
- SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
- Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
- Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models
- CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement
- EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation
- MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems
- Adversarial Attacks in Multi-Agent LLM Pipelines: Unveiling Structural Vulnerabilities in Agentic AI Architectures
- WebInject: Prompt Injection Attack to Web Agents
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems
- Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale
- Securing RAG: A Risk Assessment and Mitigation Framework
- RAG-TESTER: Automated End-to-End Testing of Retrieval-Augmented Large Language Models
- LM-Scout: Analyzing the Security of Language Model Integration in Android Apps
- GRADA: Graph-based Reranking against Adversarial Documents Attack
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
- Practical Reasoning Interruption Attacks on Reasoning Large Language Models
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents
- Dependency-Aware Privacy for Multi-turn Agents
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
- Prompt engineering for bibliographic web-scraping
- Attack and defense techniques in large language models: A survey and new perspectives
- LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
- Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware
- AgentWall: A Runtime Safety Layer for Local AI Agents
- Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
- Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
- Governable Individuals: An Identity Layer for Embodied Agents That Keep Learning
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- "Humans welcome to observe": A First Look at the Agent Social Network Moltbook
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
- Cyber Threat Intelligence for Artificial Intelligence Systems
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
- Systems-Level Attack Surface of Edge Agent Deployments on IoT
- Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
- Prompt Injection as Role Confusion
- Behavioral Integrity Verification for AI Agent Skills
- Positive Alignment: Artificial Intelligence for Human Flourishing
- Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study
- Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
- Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents
- Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
- MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
- Generative AI in Financial Institution: A Global Survey of Opportunities, Threats, and Regulation
- Traceback of Poisoning Attacks to Retrieval-Augmented Generation
- CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks
- ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- ACE: A Security Architecture for LLM-Integrated App Systems
- Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
- Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
- Prompt Injection Attack to Tool Selection in LLM Agents
- When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
- Deployment-Time Memorization in Foundation-Model Agents
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
- Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
- Owner-Harm: A Missing Threat Model for AI Agent Safety
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- Buy versus Build an LLM: A Decision Framework for Governments
- Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
- Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
- Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
- Understanding and Mitigating Risks of Generative AI in Financial Services
- RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
- ThreMoLIA: Threat Modeling of Large Language Model-Integrated Applications
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
- PIArena: A Platform for Prompt Injection Evaluation
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- SoK: Blockchain Agent-to-Agent Payments
- Test-time reasoning effort and unauthorized tool use in language-model agents: a prespecified equivalence study
- A Security-Oriented Lifecycle Model for Large Language Model Systems
- Secure and Efficient Access Control for Computer-Use Agents via Context Space
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web Agents
- PolicyGuard: Prompt-Configurable Semantic DLP for LLM Coding Agents
- "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
- Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills
- AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection
- Practical Poisoning Attacks against Retrieval-Augmented Generation
- Contextual Agentic Memory is a Memo, Not True Memory
- An Inline Control Architecture for Language Models in Intelligent Transportation Systems
- Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
- Frontier AI's Impact on the Cybersecurity Landscape
- Information Leakage of Sentence Embeddings via Generative Embedding Inversion Attacks
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models
- Breaking the Prompt Wall (I): A Real-World Case Study of Attacking ChatGPT via Lightweight Prompt Injection
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- Progent: Programmable Privilege Control for LLM Agents
- Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
- When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
- PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents
- Robust Context-Aware Detection of Malicious Instructions in Text
- StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
- You've Changed: Detecting Modification of Black-Box Large Language Models
- Defense against Prompt Injection Attacks via Mixture of Encodings
- AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
- Towards Evidence-Based Tech Hiring Pipelines
- Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators
Discussions
- Compromising LLM-integrated applications with indirect prompt injection [hn, 43 points, 20 comments]
- Novel Prompt Injection Threats to Application-Integrated Large Language Models [hn, 8 points, 2 comments]
- Getting more than what you've asked for: The Next Stage of Prompt Engineering [hn, 6 points, 1 comments]
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection [lobsters, 4 points, 0 comments]
- they cite this paper on AI agent hacking via "indirect prompt injection" [bsky, 2 points, 0 comments]
- Novel Prompt Injection Threats to Application-Integrated Large Language Models [hn, 1 points, 0 comments]
- I'm just going to leave this here... arxiv.org/abs/2302.12173 [bsky, 0 points, 0 comments]
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection https://arxiv.org/abs/2302.12173 [bsky, 0 points, 0 comments]
- - Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Ma [bsky, 0 points, 1 comments]
- "Compromising Real-WorldLLM-Integrated Applications with Indirect Prompt Injection" arxiv.org/pdf/2302.12173 [bsky, 0 points, 0 comments]
Related