LLM Social Simulations Are a Promising Research Method
2025/04/03 by Jacy Reese Anthis, Ryan Liu, Anthis, Jacy Reese +16 · 2 voices · 52 citations
Social Sciences · #Artificial Intelligence in Law
paper · pdf · doi:10.48550/arxiv.2504.02234
Abstract
Accurate and verifiable large language model (LLM) simulations of human research subjects promise an accessible data source for understanding human behavior and training new AI systems. However, results to date have been limited, and few social scientists have adopted this method. In this position paper, we argue that the promise of LLM social simulations can be achieved by addressing five tractable challenges. We ground our argument in a review of empirical comparisons between LLMs and human research subjects, commentaries on the topic, and related work. We identify promising directions, including context-rich prompting and fine-tuning with social science datasets. We believe that LLM social simulations can already be used for pilot and exploratory studies, and more widespread use may soon be possible with rapidly advancing LLM capabilities. Researchers should prioritize developing conceptual models and iterative evaluations to make the best use of new AI systems.
Citations
Cited by
- Computational Turing Test Reveals Systematic Differences Between Human and AI Language
- Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
- The threat of analytic flexibility in using large language models to simulate human data
- Challenges in Statistics: A Dozen Challenges in Causality and Causal Inference
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Agent-based simulation of online social networks and disinformation
- Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments
- Generative AI as Digital Representatives in Collective Decision-Making: A Game-Theoretical Approach
- MASim: Multilingual Agent-Based Simulation for Social Science
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models
- Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
- Can Intelligent User Interfaces Engage in Philosophical Discussions? A Longitudinal Study of Philosophers' Evolving Perceptions
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?
- Generative AI in Sociological Research: State of the Discipline
- A Criminology of Machines
- Rethinking LLM Human Simulation: When a Graph is What You Need
- Simulating and Experimenting with Social Media Mobilization Using LLM Agents
- Will Scaling Improve Social Simulation with LLMs?
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- LLM-augmented empirical game theoretic simulation for social-ecological systems
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Black Box Absorption: LLMs Undermining Innovative Ideas
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Tailored untruths: How personalisation challenges LLM safeguards
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
- Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
- Simulating Teams with LLM Agents: Interactive 2D Environments for Studying Human-AI Dynamics
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI
- Prototyping Digital Social Spaces through Metaphor-Driven Design: Translating Spatial Concepts into an Interactive Social Simulation
- PolicyPad: Collaborative Prototyping of LLM Policies
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- A large-scale evaluation of commonsense knowledge in humans and large language models
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
- Synthetic Data Generation for Screen Time and App Usage
- Inject, Fork, Compare: Defining an Interaction Vocabulary for Multi-Agent Simulation Platforms
- Programmable Cognitive Bias in Social Agents
- RecoWorld: Building Simulated Environments for Agentic Recommender Systems
- The Language of Approval: Identifying the Drivers of Positive Feedback Online
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making
- Valid Inference with Imperfect Synthetic Data
- Do Machines Think Emotionally? Cognitive Appraisal Analysis of Large Language Models
- Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
- How Exposed Are UK Jobs to Generative AI? Developing and Applying a Novel Task-Based Index
Discussions
Related