LLM-Based Social Simulations Require a Boundary
2025/06/24 by Zhisheng Wu, Wu, Zengqing, Run Peng +4 · 6 citations
Computer Science · Social Sciences · Business, Management and Accounting · #Multi-Agent Systems and Negotiation #Artificial Intelligence in Law #Business Process Modeling and Analysis
paper · pdf · doi:10.48550/arxiv.2506.19806
Abstract
This position paper argues that LLM-based social simulations require clear boundaries to make meaningful contributions to social science. While Large Language Models (LLMs) offer promising capabilities for simulating human behavior, their tendency to produce homogeneous outputs, acting as an "average persona", fundamentally limits their ability to capture the behavioral diversity essential for complex social dynamics. We examine why heterogeneity matters for social simulations and how current LLMs fall short, analyzing the relationship between mean alignment and variance in LLM-generated behaviors. Through a systematic review of representative studies, we find that validation practices often fail to match the heterogeneity requirements of research questions: while most papers include ground truth comparisons, fewer than half explicitly assess behavioral variance, and most that do report lower variance than human populations. We propose that researchers should: (1) match validation depth to the heterogeneity demands of their research questions, (2) explicitly report variance alongside mean alignment, and (3) constrain claims to collective-level qualitative patterns when variance is insufficient. Rather than dismissing LLM-based simulation, we advocate for a boundary-aware approach that ensures these methods contribute genuine insights to social science.
Citations
- Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics
- Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- Population-Aligned Persona Generation for LLM-based Social Simulation
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Can Generative AI agents behave like humans? Evidence from laboratory market experiments
- Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
- MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation
- Can Large Language Models Trade? Testing Financial Theories with LLM Agents in Market Simulations
- SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- LLM Social Simulations Are a Promising Research Method
- From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment
- LLM Generated Persona is a Promise with a Catch
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
- From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
- Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs: Evidence and Mechanisms for AI Agents to Inform Policy Decisions
- Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
- LLM-based Human Simulations Have Not Yet Been Reliable
- Enhancing Patient-Centric Communication: Leveraging LLMs to Simulate Patient Perspectives
- On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
- Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
- Observing Micromotives and Macrobehavior of Large Language Models
- Simulating Human-like Daily Activities with Desire-driven Autonomy
- From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents
- OASIS: Open Agent Social Interaction Simulations with One Million Agents
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
- Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
- Emergent social conventions and collective bias in LLM populations
- Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
- GenSim: A General Social Simulation Platform with Large Language Model based Agents
- Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
- POSIX: A Prompt Sensitivity Index For Large Language Models
- Decoding Echo Chambers: LLM-Powered Simulations Revealing Polarization in Social Networks
- LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
- Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
- DiverseDialogue: A Methodology for Designing Chatbots with Human-Like Diversity
- United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
- Behavioral and Topological Heterogeneities in Network Versions of Schelling's Segregation Model
- Identifying and Mitigating Social Bias Knowledge in Language Models
- Evaluating and Enhancing LLMs Agent based on Theory of Mind in Guandan: A Multi-Player Cooperative Game under Imperfect Information
- The Sociolinguistic Foundations of Language Modeling
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Simulating Classroom Education with LLM-Empowered Agents
- Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?
- Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
- Evaluating Large Language Model Biases in Persona-Steered Generation
- Knowing What Not to Do: Leverage Language Model Insights for Action Space Pruning in Multi-agent Reinforcement Learning
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models
- Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
- Strategic Interactions between Large Language Models-based Agents in Beauty Contests
- Measuring Social Norms of Large Language Models
- Explaining Large Language Models Decisions Using Shapley Values
- Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
- Can Large Language Models Play Games? A Case Study of A Self-Play Approach
- Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
- Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation
- Are Large Language Models (LLMs) Good Social Predictors?
- Shall We Team Up: Exploring Spontaneous Cooperation of Competing LLM Agents
- Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
- K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
- Large language models that replace human participants can harmfully misportray and flatten identity groups
- Computational Experiments Meet Large Language Model Based Agents: A Survey and Perspective
- Rethinking Interpretability in the Era of Large Language Models
- Prompt Smells: An Omen for Undesirable Generative AI Outputs
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- Large Language Models Empowered Agent-based Modeling and Simulation: A Survey and Perspectives
- War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
- Simulating Opinion Dynamics with Networks of LLM-based Agents
- Multiagent Simulators for Social Networks
- When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
- Smart Agent-Based Modeling: On the Use of Large Language Models in Computer Simulations
- Do LLMs exhibit human-like response biases? A case study in survey design
- Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
- CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations
- EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities
- Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View
- LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
- Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- "Guinea Pig Trials" Utilizing GPT: A Novel Smart Agent-Based Modeling Approach for Studying Firm Competition and Collusion
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
- AgentBench: Evaluating LLMs as Agents
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- S3: Social-network Simulation System with Large Language Model-Empowered Agents
- ChatDev: Communicative Agents for Software Development
- Personality Traits in Large Language Models
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- User Behavior Simulation with Large Language Model based Agents
- Strategic Reasoning with Language Models
- Playing repeated games with large language models
- Can Large Language Models Transform Computational Social Science?
- Can Large Language Models Transform Computational Social Science?
- Generative Agents: Interactive Simulacra of Human Behavior
- Whose Opinions Do Language Models Reflect?
- Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?
- Out of One, Many: Using Language Models to Simulate Human Samples
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Social Simulacra: Creating Populated Prototypes for Social Computing Systems
- What is a social pattern? Rethinking a central social science term
- Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
- Should social science be more solution-oriented?
- Estimating the reproducibility of psychological science
- Social Simulation in the Social Sciences
- Statistical physics of social dynamics
- Anomalous fluctuations in Minority Games and related multi-agent models of financial markets
- Threshold Models of Collective Behavior
- Dynamic models of segregation†
- Simulating Human Strategic Behavior: Comparing Single and Multi-agent LLMs
- Language Model Alignment in Multilingual Trolley Problems
Cited by
Related