Why Do Multi-Agent LLM Systems Fail?
2025/03/17 by Cemri, Mert, Pan, Melissa Z., Yang, Shuyi +10 · 9 voices · 117 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2503.13657
Abstract
Despite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal. This gap highlights a critical need for a principled understanding of why MAS fail. Addressing this question requires systematic identification and analysis of failure patterns. We introduce MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks. MAST-Data is the first multi-agent system dataset to outline the failure dynamics in MAS for guiding the development of better future systems. To enable systematic classification of failures for MAST-Data, we build the first Multi-Agent System Failure Taxonomy (MAST). We develop MAST through rigorous analysis of 150 traces, guided closely by expert human annotators and validated by high inter-annotator agreement (kappa = 0.88). This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification. To enable scalable annotation, we develop an LLM-as-a-Judge pipeline with high agreement with human annotations. We leverage MAST and MAST-Data to analyze failure patterns across models (GPT4, Claude 3, Qwen2.5, CodeLlama) and tasks (coding, math, general agent), demonstrating improvement headrooms from better MAS design. Our analysis provides insights revealing that identified failures require more sophisticated solutions, highlighting a clear roadmap for future research. We publicly release our comprehensive dataset (MAST-Data), the MAST, and our LLM annotator to facilitate widespread research and development in MAS.
Cited by
- Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
- Towards an Agent Operating System - Lessons from Classical and Cloud OS
- Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
- Does It Tie Out? Towards Autonomous Legal Agents in Venture Capital
- Specification and Detection of LLM Code Smells
- Let the Barbarians In: How AI Can Accelerate Systems Performance Research
- Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
- SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
- Evolving Excellence: Automated Optimization of LLM-based Agents
- Insured Agents: A Decentralized Trust Insurance Mechanism for Agentic Economy
- Towards a Science of Scaling Agent Systems
- Single-Agent Scaling Fails Multi-Agent Intelligence: Towards Foundation Models with Native Multi-Agent Intelligence
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- LLM Harms: A Taxonomy and Discussion
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Measuring Agents in Production
- Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
- Process-Centric Analysis of Agentic Software Systems
- How Far Are We from Genuinely Useful Deep Research Agents?
- An Empirical Study of Agent Developer Practices in AI Agent Frameworks
- Crystalyse: a multi-tool agent for materials design
- FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections
- SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
- Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
- Latent Collaboration in Multi-Agent Systems
- Fara-7B: An Efficient Agentic Model for Computer Use
- AnimAgents: Coordinating Multi-Stage Animation Pre-Production with Human-Multi-Agent Collaboration
- Sensorium Arc: AI Agent System for Oceanic Data Exploration and Interactive Eco-Art
- Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming
- Multi-Agent Collaborative Fuzzing with Continuous Reflection for Smart Contracts Vulnerability Detection
- Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
- Convergence dynamics of Agent-to-Agent Interactions with Misaligned objectives
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- Detecting Silent Failures in Multi-Agentic AI Trajectories
- GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
- Sherlock: Reliable and Efficient Agentic Workflow Execution
- Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
- Resolving Java Code Repository Issues with iSWE Agent
- Bridging LLM Planning Agents and Formal Methods: A Case Study in Plan Verification
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- CoDA: Agentic Systems for Collaborative Data Visualization
- The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- CooperBench: Why Coding Agents Cannot be Your Teammates Yet
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Code Aesthetics with Agentic Reward Feedback
- Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
- Thought Communication in Multiagent Collaboration
- Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
- Automating Structural Engineering Workflows with Large Language Model Agents
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search
- Effective Strategies for Asynchronous Software Engineering Agents
- BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
- How can we assess human-agent interactions? Case studies in software agent design
- Co-TAP: Three-Layer Agent Interaction Protocol Technical Report
- AgentAsk: Multi-Agent Systems Need to Ask
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization
- Barbarians at the Gate: How AI is Upending Systems Research
- Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning
- On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
- CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage
- Where LLM Agents Fail and How They can Learn From Failures
- Agentic Services Computing
- MAS2: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems
- LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research
- Demystifying the Lifecycle of Failures in Platform-Orchestrated Agentic Workflows
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems
- BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
- Evaluating LLM Generated Detection Rules in Cybersecurity
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
- Who is Introducing the Failure? Automatically Attributing Failures of Multi-Agent Systems via Spectrum Analysis
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- When Agents go Astray: Course-Correcting SWE Agents with PRMs
- ProST: Progressive Sub-task Training for Pareto-Optimal Multi-agent Systems Using Small Language Models
- Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- COCO: Cognitive Operating System with Continuous Oversight for Multi-Agent Workflow Reliability
- LumiMAS: A Comprehensive Framework for Real-Time Monitoring and Enhanced Observability in Multi-Agent Systems
- Benchmarking LLM-based Agents for Single-cell Omics Analysis
- SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows
- Multi-agent systems for chemical engineering: A review and perspective
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- From MAS to MARS: Coordination Failures and Reasoning Trade-offs in Hierarchical Multi-Agent Robotic Systems within a Healthcare Scenario
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Balancing Information Accuracy and Response Timeliness in Networked LLMs
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games
- Large Language Model-based Data Science Agent: A Survey
- A Survey on Agent Workflow -- Status and Future
Discussions
- Why do multi-agent LLM systems fail? arxiv.org/abs/2503.13657 #AI #MLsky #llms [bsky, 4 points, 0 comments]
- arxiv.org/abs/2503.13657
New study identifies 14 failure modes in multi-agent LLM systems across 150+ tasks. Despite the hype, multi-agent systems show minimal performance gains vs single agents. Fai [bsky, 4 points, 0 comments]
- arxiv.org/pdf/2503.13657 Why do multi-agent LLM systems fail. #AI #airesearch [bsky, 3 points, 0 comments]
- “Happy families are all alike; each unhappy family is unhappy in its own way.”
Researchers from @ucberkeleyofficial.bsky.social quote Tolstoy to explain how multi-agent LLM systems fail (as much as 7 [bsky, 3 points, 0 comments]
- Many have been excited about Multi-Agent Systems where multiple AI agents work together to solve problems. Despite the hype, these systems often fail to outperform single-agent approaches on standard [bsky, 2 points, 0 comments]
- Why Do Multi-Agent LLM Systems Fail? [hn, 1 points, 0 comments]
- Why Do Multi-Agent LLM Systems Fail? [hn, 1 points, 0 comments]
- 3/n - They showed small interventions (like role clarification and verification steps) led to 9–16% improvements (picture). MAST moves MAS development closer to science and engineering rather than gue [bsky, 0 points, 0 comments]
- Counter-evidence, though. Berkeley's MAST paper clocks 41 to 86 percent failure rates across 1,600+ production traces; Walden Yan at Cognition argues sub-agents drift because they share messages, not [bsky, 0 points, 1 comments]
Related