Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
2025/10/15 by Hu, Jinwei, Dong, Yi, Ao, Shuang +6 · 2 citations
#FOS: Computer and information sciences #Multiagent Systems (cs.MA)
paper · doi:10.48550/arxiv.2510.14008
Abstract
LLM-powered Multi-Agent Systems (LLM-MAS) unlock new potentials in distributed reasoning, collaboration, and task generalization but also introduce additional risks due to unguaranteed agreement, cascading uncertainty, and adversarial vulnerabilities. We argue that ensuring responsible behavior in such systems requires a paradigm shift: from local, superficial agent-level alignment to global, systemic agreement. We conceptualize responsibility not as a static constraint but as a lifecycle-wide property encompassing agreement, uncertainty, and security, each requiring the complementary integration of subjective human-centered values and objective verifiability. Furthermore, a dual-perspective governance framework that combines interdisciplinary design with human-AI collaborative oversight is essential for tracing and ensuring responsibility throughout the lifecycle of LLM-MAS. Our position views LLM-MAS not as loose collections of agents, but as unified, dynamic socio-technical systems that demand principled mechanisms to support each dimension of responsibility and enable ethically aligned, verifiably coherent, and resilient behavior for sustained, system-wide agreement.
Citations
- Tapas Are Free! Training-Free Adaptation of Programmatic Agents via LLM-Guided Program Synthesis in Dynamic Environments
- Hierarchical Testing with Rabbit Optimization for Industrial Cyber-Physical Systems
- Enhancing Robustness of LLM-Driven Multi-Agent Systems through Randomized Smoothing
- TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
- AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society
- Large Language Model Based Multi-Agent System Augmented Complex Event Processing Pipeline for Internet of Multimedia Things
- A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions
- LLM-PySC2: Starcraft II learning environment for Large Language Models
- An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-Making
- JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
- Moral Alignment for LLM Agents
- Control Industrial Automation System with Large Language Model Agents
- Strategic Collusion of LLM Agents: Market Division in Multi-Commodity Competitions
- Understanding Knowledge Drift in LLMs through Misinformation
- SciAgents: Automating scientific discovery through multi-agent intelligent graph reasoning
- Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
- Trust-Oriented Adaptive Guardrails for Large Language Models
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
- Uncertainty is Fragile: Manipulating Uncertainty in Large Language Models
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
- SAIL: Self-Improving Efficient Online Alignment of Large Language Models
- Autonomous Agents for Collaborative Task under Information Asymmetry
- Safeguarding Large Language Models: A Survey
- Evaluating Uncertainty-based Failure Detection for Closed-Loop LLM Planners
- Exploring Prosocial Irrationality for LLM Agents: A Social Cognition View
- Aligning LLM Agents by Learning Latent Preference from User Edits
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
- AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration
- Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
- Advancing LLM Reasoning Generalists with Preference Trees
- LUQ: Long-text Uncertainty Quantification for LLMs
- Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
- Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
- LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
- Probabilistically Correct Language-based Multi-Robot Planning using Conformal Prediction
- Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents
- Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis
- Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
- In-context learning agents are asymmetric belief updaters
- LLM Multi-Agent Systems: Challenges and Open Problems
- A Multi-Agent Conversational Recommender System
- Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning
- Security and Privacy Challenges of Large Language Models: A Survey
- Towards Uncertainty-Aware Language Agent
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Self-Rewarding Language Models
- Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects
- The Persuasive Power of Large Language Models
- Large Language Models Empowered Agent-based Modeling and Simulation: A Survey and Perspectives
- The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
- Scalable AI Safety via Doubly-Efficient Debate
- MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- LLM4Drive: A Survey of Large Language Models for Autonomous Driving
- Multi-Agent Consensus Seeking via Large Language Models
- The Consensus Game: Language Model Generation via Equilibrium Search
- In-Context Unlearning: Language Models as Few Shot Unlearners
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
- Improving the Reliability of Large Language Models by Leveraging Uncertainty-Aware In-Context Learning
- Balancing Autonomy and Alignment: A Multi-Dimensional Taxonomy for Autonomous LLM-powered Multi-Agent Architectures
- Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Large Language Model Alignment: A Survey
- Can LLM-Generated Misinformation Be Detected?
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- Quantifying the Impact of Large Language Models on Collective Opinion Dynamics
- Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
- Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models
- What, Indeed, is an Achievable Provable Guarantee for Learning-Enabled Safety Critical Systems
- AlpaGasus: Training A Better Alpaca with Fewer Data
- ChatDev: Communicative Agents for Software Development
- Self-Adaptive Large Language Model (LLM)-Based Multiagent Systems
- Epidemic Modeling with Generative Agents
- Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners
- Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
- Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
- A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
- WizardLM: Empowering large pre-trained language models to follow complex instructions
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
- MoralDial: A Framework to Train and Evaluate Moral Dialogue Systems via Moral Discussions
- Large Language Models Can Self-Improve
- Improving alignment of dialogue agents via targeted human judgements
- Symbolic Runtime Verification for Monitoring under Uncertainties and Assumptions
- Language Models (Mostly) Know What They Know
- Trust-based Consensus in Multi-Agent Reinforcement Learning Systems
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Training language models to follow instructions with human feedback
- Survey of Hallucination in Natural Language Generation
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Truthful AI: Developing and governing AI that does not lie
- Conformal Bayesian Computation
- Preference learning along multiple criteria: A game-theoretic perspective
- Neural-Symbolic Integration: A Compositional Perspective
- Learning to summarize from human feedback
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Fine-Tuning Language Models from Human Preferences
- Supervising strong learners by amplifying weak experts
- AI safety via debate
- Proximal Policy Optimization Algorithms
- Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks
- Introspective Planning: Aligning Robots' Uncertainty with Inherent Task Ambiguity
- DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics
- Large Model Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends
Cited by
Related