Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
2025/09/25 by He, Ping, Li, Changjiang, Zhao, Binbin +2 · 3 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Software Engineering (cs.SE)
paper · doi:10.48550/arxiv.2509.21011
Abstract
The remarkable capability of large language models (LLMs) has led to the wide application of LLM-based agents in various domains. To standardize interactions between LLM-based agents and their environments, model context protocol (MCP) tools have become the de facto standard and are now widely integrated into these agents. However, the incorporation of MCP tools introduces the risk of tool poisoning attacks, which can manipulate the behavior of LLM-based agents. Although previous studies have identified such vulnerabilities, their red teaming approaches have largely remained at the proof-of-concept stage, leaving the automatic and systematic red teaming of LLM-based agents under the MCP tool poisoning paradigm an open question. To bridge this gap, we propose AutoMalTool, an automated red teaming framework for LLM-based agents by generating malicious MCP tools. Our extensive evaluation shows that AutoMalTool effectively generates malicious MCP tools capable of manipulating the behavior of mainstream LLM-based agents while evading current detection mechanisms, thereby revealing new security risks in these agents.
Citations
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- WebSailor: Navigating Super-human Reasoning for Web Agent
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- Deep Research Agents: A Systematic Examination And Roadmap
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Design Patterns for Securing LLM Agents against Prompt Injections
- A Critical Evaluation of Defenses against Prompt Injection Attacks
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
- Prompt Injection Attack to Tool Selection in LLM Agents
- DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
- Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- Defeating Prompt Injections by Design
- Defending against Adversarial Malware Attacks on ML-based Android Malware Detection Systems
- From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- A Survey on LLM-as-a-Judge
- Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
- RedCode: Risky Code Execution and Generation Benchmark for Code Agents
- Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- FinRobot: An Open-Source AI Agent Platform for Financial Applications using Large Language Models
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
- Optimization-based Prompt Injection Attack to LLM-as-a-Judge
- Automatic and Universal Prompt Injection Attacks against Large Language Models
- A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist
- StruQ: Defending Against Prompt Injection with Structured Queries
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- Efficient Query-Based Attack against ML-Based Android Malware Detection under Zero Knowledge Setting
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
- Mind2Web: Towards a Generalist Agent for the Web
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Prompt Injection attack against LLM-integrated Applications
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
- Taxonomy of Attacks on Open-Source Software Supply Chains
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Matching Networks for One Shot Learning
- Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
- Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents
- PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification
Cited by
Related