Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
2025/10/15 by Chen, Jinkun, Badshah, Sher, Yu, Xuemin +1
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multiagent Systems (cs.MA)
paper · doi:10.48550/arxiv.2510.13982
Abstract
What if artificial agents could not just communicate, but also evolve, adapt, and reshape their worlds in ways we cannot fully predict? With llm now powering multi-agent systems and social simulations, we are witnessing new possibilities for modeling open-ended, ever-changing environments. Yet, most current simulations remain constrained within static sandboxes, characterized by predefined tasks, limited dynamics, and rigid evaluation criteria. These limitations prevent them from capturing the complexity of real-world societies. In this paper, we argue that static, task-specific benchmarks are fundamentally inadequate and must be rethought. We critically review emerging architectures that blend llm with multi-agent dynamics, highlight key hurdles such as balancing stability and diversity, evaluating unexpected behaviors, and scaling to greater complexity, and introduce a fresh taxonomy for this rapidly evolving field. Finally, we present a research roadmap centered on open-endedness, continuous co-evolution, and the development of resilient, socially aligned AI ecosystems. We call on the community to move beyond static paradigms and help shape the next generation of adaptive, socially-aware multi-agent simulations.
Citations
- TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- Survey on Evaluation of LLM-based Agents
- TwinMarket: A Scalable Behavioral and Social Simulation for Financial Markets
- Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
- GAI: Generative Agents for Innovation
- A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application
- From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
- Incentives to Build Houses, Trade Houses, or Trade House Building Skills in Simulated Worlds under Various Governing Systems or Institutions: Comparing Multi-agent Reinforcement Learning to Generative Agent-based Model
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
- Thanos: Enhancing Conversational Agents with Skill-of-Mind-Infused Large Language Model
- OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
- Towards Trustworthy Knowledge Graph Reasoning: An Uncertainty Aware Perspective
- Causal Representation Learning with Generative Artificial Intelligence: Application to Texts as Treatments
- StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
- LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
- "A Woman is More Culturally Knowledgeable than A Man?": The Effect of Personas on Cultural Norm Interpretation in LLMs
- Model Tells Itself Where to Attend: Faithfulness Meets Automatic Attention Steering
- Good Idea or Not, Representation of LLM Could Tell
- Generative Agent-Based Models for Complex Systems Research: a review
- From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
- Mimicking the Mavens: Agent-based Opinion Synthesis and Emotion Prediction for Social Media Influencers
- Recursive Introspection: Teaching Language Model Agents How to Self-Improve
- Prover-Verifier Games improve legibility of LLM outputs
- Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
- Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy
- Using Grammar Masking to Ensure Syntactic Validity in LLM-based Modeling Tasks
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
- Venn Diagram Prompting : Accelerating Comprehension with Scaffolding Effect
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
- Navigating LLM Ethics: Advancements, Challenges, and Future Directions
- Exploring the Potential of Conversational AI Support for Agent-Based Social Simulation Model Design
- Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- Unifying Large Language Model and Deep Reinforcement Learning for Human-in-Loop Interactive Socially-aware Navigation
- TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
- Affordable Generative Agents
- Human-mediated Large Language Models for Robotic Intervention in Children with Autism Spectrum Disorders
- What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection
- Weak-to-Strong Jailbreaking on Large Language Models
- "You tell me": A Dataset of GPT-4-Based Behaviour Change Support Conversations
- Large language model empowered participatory urban planning
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Metacognition is all you need? Using Introspection in Generative Agents to Improve Goal-directed Behavior
- Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects
- Viz: A QLoRA-based Copyright Marketplace for Legally Compliant Generative AI
- Generative agents in the streets: Exploring the use of Large Language Models (LLMs) in collecting urban perceptions
- Simulating Public Administration Crisis: A Novel Generative Agent-Based Simulation System to Lower Technology Barriers in Social Science Research
- Identifying and Mitigating Vulnerabilities in LLM-Integrated Applications
- PromptMix: A Class Boundary Augmentation Method for Large Language Model Distillation
- Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
- Character-LLM: A Trainable Agent for Role-Playing
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking Partner
- Balancing Autonomy and Alignment: A Multi-Dimensional Taxonomy for Autonomous LLM-powered Multi-Agent Architectures
- Lyfe Agents: Generative agents for low-cost real-time social interactions
- Chatmap : Large Language Model Interaction with Cartographic Data
- SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking Data
- TradingGPT: Multi-Agent System with Layered Memory and Distinct Characters for Enhanced Financial Trading Performance
- MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records
- AgentSims: An Open-Source Sandbox for Large Language Model Evaluation
- AgentBench: Evaluating LLMs as Agents
- Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition
- Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
- Wireless Multi-Agent Generative AI: From Connected Intelligence to Collective Intelligence
- ToolQA: A Dataset for LLM Question Answering with External Tools
- Trapping LLM Hallucinations Using Tagged Context Prompts
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- What In-Context Learning "Learns" In-Context: Disentangling Task Recognition and Task Learning
- T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering
- The Role of Summarization in Generative Agents: A Preliminary Perspective
- The Internal State of an LLM Knows When It's Lying
- nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales
- Generative Agents: Interactive Simulacra of Human Behavior
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Reflexion: Language Agents with Verbal Reinforcement Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
- Language Models are Few-Shot Learners
- Relational Forward Models for Multi-Agent Learning
- Emergence of Grounded Compositional Language in Multi-Agent Populations
- Multi-Agent Cooperation and the Emergence of (Natural) Language
Related