Large language models that replace human participants can harmfully misportray and flatten identity groups
2024/02/02 by Angelina Wang, Jamie Morgenstern, Wang, Angelina +3 · 7 voices · 78 citations
Computer Science · Psychology · #Aesthetics #Art #Computer science #Identity (music) #Linguistics #Natural Language Processing Techniques #Natural language processing #Philosophy #Psychology #Topic Modeling #cs.CY
paper · pdf · doi:10.48550/arxiv.2402.01908
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/02 · openalex created_date 2024/02/07 · openalex updated_date 2026/08/05
Abstract
Large language models (LLMs) are increasing in capability and popularity, propelling their application in new domains -- including as replacements for human participants in computational social science, user testing, annotation tasks, and more. In many settings, researchers seek to distribute their surveys to a sample of participants that are representative of the underlying human population of interest. This means in order to be a suitable replacement, LLMs will need to be able to capture the influence of positionality (i.e., relevance of social identities like gender and race). However, we show that there are two inherent limitations in the way current LLMs are trained that prevent this. We argue analytically for why LLMs are likely to both misportray and flatten the representations of demographic groups, then empirically show this on 4 LLMs through a series of human studies with 3200 participants across 16 demographic identities. We also discuss a third limitation about how identity prompts can essentialize identities. Throughout, we connect each limitation to a pernicious history of epistemic injustice against the value of lived experiences that explains why replacement is harmful for marginalized demographic groups. Overall, we urge caution in use cases where LLMs are intended to replace human participants whose identities are relevant to the task at hand. At the same time, in cases where the benefits of LLM replacement are determined to outweigh the harms (e.g., the goal is to supplement rather than fully replace, engaging human participants may cause them harm), we provide inference-time techniques that we empirically demonstrate do reduce, but do not remove, these harms.
Cited by
- Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
- Co-design of LLM-based preference agents: participation may drive overtrust
- Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling
- More Is Not More: What Matters for Diversity in LLM Opinions?
- Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences
- The threat of analytic flexibility in using large language models to simulate human data
- Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- Correlated Errors in Large Language Models
- A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
- Value Profiles for Encoding Human Variation
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Personalization, Personas, and Forecasting in Value Alignment
- Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments
- Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
- Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors
- Social Perceptions of English Spelling Variation on Twitter: A Comparative Analysis of Human and LLM Responses
- Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?
- German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys
- DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization
- Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback
- Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
- Effectiveness of Large Language Models in Simulating Regional Psychological Structures: An Empirical Examination of Personality and Subjective Well-being
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Transphobia Is in the Eye of the Prompter: Trans-Centered Perspectives on Large Language Models
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Qualitative Research in an Era of AI: A Pragmatic Approach to Data Analysis, Workflow, and Computation
- Finetuning LLMs for Human Behavior Prediction in Social Science Experiments
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- On the Alignment of Large Language Models with Global Human Opinion
- Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants
- AIM-Bench: Evaluating Decision-making Biases of Agentic LLM as Inventory Manager
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models
- Valuing Time in Silicon: Can Large Language Models Replicate Human Value of Travel Time
- Exploring LLMs for Automated Pre-Testing of Cross-Cultural Surveys
- Improving the Distributional Alignment of LLMs using Supervision
- How large language models judge and influence human cooperation
- LLM-Based Social Simulations Require a Boundary
- Large Language Models as Psychological Simulators: A Methodological Guide
- From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
- Exploring MLLMs Perception of Network Visualization Principles
- Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
- AI Agent Behavioral Science
- Beyond Static Responses: Multi-Agent LLM Systems as a New Paradigm for Social Science Research
- ValueSim: Generating Backstories to Model Individual Value Systems
- Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
- Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models
- Persona Alchemy: Designing, Evaluating, and Implementing Psychologically-Grounded LLM Agents for Diverse Stakeholder Representation
- Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
- Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation
- Multilingual Prompting for Improving LLM Generation Diversity
- MindVote: When AI Meets the Wild West of Social Media Opinion
- Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays
- What LLMs Think When You Don't Tell Them What to Think About?
- Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
- LLM Consumer Behavior Theory: Foundations of a Novel Research Field
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- Aspirational Affordances of AI
- Mind the Gaps: Mixture-of-Minds for Human Simulation
- Should you use LLMs to simulate opinions? Quality checks for early-stage deliberation
- Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
- Cultural Learning-Based Culture Adaptation of Language Models
Discussions
- LLMs replacing human participants harmfully misportray, flatten identity groups [hn, 19 points, 9 comments]
- At least people are taking them to task. I have a paper coming out on this too soon. But it's not a marginal view that this is okay sadly. arxiv.org/abs/2402.019... [bsky, 12 points, 1 comments]
- Open-access version of this paper here....
arxiv.org/pdf/2402.01908
"...we urge caution in use cases where LLMs are intended to replace human participants whose identities are relevant to the task a [bsky, 11 points, 2 comments]
- on this point
arxiv.org/abs/2402.01908 [bsky, 11 points, 0 comments]
- Yes! The preprint is available here: arxiv.org/abs/2402.01908 [bsky, 2 points, 1 comments]
- related paper: "Large language models should not replace human participants because they can misportray and flatten identity groups"
arxiv.org/abs/2402.01908 [bsky, 1 points, 0 comments]
- LLMs replacing human participants harmfully misportray, flatten identity groups https:// arxiv.org/abs/2402.01908 # arxiv # llm # llms [mastodon, 0 points, 0 comments]
Related