Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
2025/10/27 by Datta, Shrestha, Nahin, Shahriar Kabir, Chhabra, Anshuman +1 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.23883
Abstract
Agentic AI systems powered by large language models (LLMs) and endowed with planning, tool use, memory, and autonomy, are emerging as powerful, flexible platforms for automation. Their ability to autonomously execute tasks across web, software, and physical environments creates new and amplified security risks, distinct from both traditional AI safety and conventional software security. This survey outlines a taxonomy of threats specific to agentic AI, reviews recent benchmarks and evaluation methodologies, and discusses defense strategies from both technical and governance perspectives. We synthesize current research and highlight open challenges, aiming to support the development of secure-by-design agent systems.
Citations
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
- Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
- Evaluation and Benchmarking of LLM Agents: A Survey
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree
- Prompt Injection 2.0: Hybrid AI Threats
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
- Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
- MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
- OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- Design Patterns for Securing LLM Agents against Prompt Injections
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
- Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
- Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
- LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
- Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
- Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
- Generative to Agentic AI: Survey, Conceptualization, and Challenges
- Towards a HIPAA Compliant Agentic AI System in Healthcare
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
- AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
- DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agents
- Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
- Agents Under Siege: Breaking Pragmatic Multi-Agent LLM Systems with Optimized Prompt Attacks
- SandboxEval: Towards Securing Test Environment for Untrusted Code
- sudo rm -rf agenticsecurity
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- Defeating Prompt Injections by Design
- Why Do Multi-Agent LLM Systems Fail?
- Prompt Injection Detection and Mitigation via AI Multi-Agent NLP Frameworks
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
- SafeArena: Evaluating the Safety of Autonomous Web Agents
- BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
- Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
- Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems
- Can Indirect Prompt Injection Attacks Be Detected and Removed?
- Red-Teaming LLM Multi-Agent Systems via Communication Attacks
- A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
- Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks
- Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
- Fully Autonomous AI Agents Should Not be Developed
- s1: Simple test-time scaling
- Unraveling Indirect In-Context Learning Using Influence Functions
- The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
- SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
- The BrowserGym Ecosystem for Web Agent Research
- Examining Identity Drift in Conversations of LLM Agents
- A Survey on LLM-as-a-Judge
- Attacking Vision-Language Computer Agents via Pop-ups
- Attention Tracker: Detecting Prompt Injection Attacks in LLMs
- AdvAgent: Controllable Blackbox Red-teaming on Web Agents
- Imprompter: Tricking LLM Agents into Improper Tool Use
- Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
- AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
- Agent-as-a-Judge: Evaluate Agents with Agents
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- ToolGen: Unified Tool Retrieval and Calling via Generation
- Permissive Information-Flow Analysis for Large Language Models
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
- CLIBE: Detecting Dynamic Backdoors in Transformer-based NLP Models
- PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
- HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model
- Can We Rely on LLM Agents to Draft Long-Horizon Plans? Let's Take TravelPlanner as an Example
- MPC-Minimized Secure LLM Inference
- On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
- Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
- The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- GTA: A Benchmark for General Tool Agents
- R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
- AI Agents That Matter
- CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
- AgentDojo-PROV: A W3C PROV-O Corpus of LLM Agent Executions
- Dissecting Adversarial Robustness of Multimodal LM Agents
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Current state of LLM Risks and AI Guardrails
- OpenVLA: An Open-Source Vision-Language-Action Model
- Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
- Improving Alignment and Robustness with Circuit Breakers
- Get my drift? Catching LLM Task Drift with Activation Deltas
- A Survey of Multimodal Large Language Model from A Data-centric Perspective
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
- Human-Imperceptible Retrieval Poisoning Attacks in LLM-Powered Applications
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Group-Aware Coordination Graph for Multi-Agent Reinforcement Learning
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- LLM Agents can Autonomously Exploit One-day Vulnerabilities
- Empowering Biomedical Discovery with AI Agents
- Long-context LLMs Struggle with Long In-context Learning
- Large Language Model Evaluation Via Multi AI Agents: Preliminary results
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Automatic and Universal Prompt Injection Attacks against Large Language Models
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- Learning to Use Tools via Cooperative and Interactive Agents
- A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
- Watermarking Makes Language Models Radioactive
- Aligning Individual and Collective Objectives in Multi-Agent Cooperation
- Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- StruQ: Defending Against Prompt Injection with Structured Queries
- LLM Agents can Autonomously Hack Websites
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- An Early Categorization of Prompt Injection Attacks on Large Language Models
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
- Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Jatmo: Prompt Injection Defense by Task-Specific Finetuning
- The Adaptive Arms Race: Redefining Robustness in AI Security
- HuRef: HUman-REadable Fingerprint for Large Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Poisoning Retrieval Corpora by Injecting Adversarial Passages
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
- AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
- AutoAgents: A Framework for Automatic Agent Generation
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
- Privacy Side Channels in Machine Learning Systems
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- Detecting Language Model Attacks with Perplexity
- Adversarial Illusions in Multi-Modal Embeddings
- AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
- AgentBench: Evaluating LLMs as Agents
- From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application?
- Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Autonomous Tester Agent Benchmark
- Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- A Comprehensive Overview of Large Language Models
- Tools for Verifying Neural Models' Training Data
- Mind2Web: Towards a Generalist Agent for the Web
- From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Multimodal Web Navigation with Instruction-Finetuned Foundation Models
- Otter: A Multi-Modal Model with In-Context Instruction Tuning
- Generative Agents: Interactive Simulacra of Human Behavior
- GPT-4 Technical Report
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Ignore Previous Prompt: Attack Techniques For Language Models
- DP-Rewrite: Towards Reproducibility and Transparency in Differentially Private Text Rewriting
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- Are Large Pre-Trained Language Models Leaking Your Personal Information?
- Just Fine-tune Twice: Selective Differential Privacy for Large Language Models
- Training language models to follow instructions with human feedback
- Application of Homomorphic Encryption in Medical Imaging
- BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning
- Evaluating Large Language Models Trained on Code
- Hidden Backdoors in Human-Centric Language Models
- Data Poisoning Attacks Against Federated Learning Systems
- Language Models are Few-Shot Learners
- Adaptive Reward-Poisoning Attacks against Reinforcement Learning
- Heuristic Black-box Adversarial Attacks on Video Recognition Models
- Online Synthesis for Runtime Enforcement of Safety in Multi-Agent Systems
- Adversarial Objects Against LiDAR-Based Autonomous Driving Systems
- Fooling automated surveillance cameras: adversarial patches to attack\n person detection
- Black-box Adversarial Attacks on Video Recognition Models
- Transferable Adversarial Attacks for Image and Video Object Detection
- Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods
- Understanding Black-box Predictions via Influence Functions
- Intriguing properties of neural networks
- The epistemology of a rule-based expert system —a framework for explanation
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- TrustAgent: Towards Safe and Trustworthy LLM-based Agents
Cited by
Related