Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
2025/07/02 by Mannekote, Amogh, Davies, Adam, Li, Guohao +4
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2507.02197
Abstract
As LLMs are increasingly studied as role-playing agents to generate synthetic data for human behavioral research, ensuring that their outputs remain coherent with their assigned roles has become a critical concern. In this paper, we investigate how consistently LLM-based role-playing agents' stated beliefs about the behavior of the people they are asked to role-play ("what they say") correspond to their actual behavior during role-play ("how they act"). Specifically, we establish an evaluation framework to rigorously measure how well beliefs obtained by prompting the model can predict simulation outcomes in advance. Using an augmented version of the GenAgents persona bank and the Trust Game (a standard economic game used to quantify players' trust and reciprocity), we introduce a belief-behavior consistency metric to systematically investigate how it is affected by factors such as: (1) the types of beliefs we elicit from LLMs, like expected outcomes of simulations versus task-relevant attributes of individual characters LLMs are asked to simulate; (2) when and how we present LLMs with relevant information about Trust Game; and (3) how far into the future we ask the model to forecast its actions. We also explore how feasible it is to impose a researcher's own theoretical priors in the event that the originally elicited beliefs are misaligned with research objectives. Our results reveal systematic inconsistencies between LLMs' stated (or imposed) beliefs and the outcomes of their role-playing simulation, at both an individual- and population-level. Specifically, we find that, even when models appear to encode plausible beliefs, they may fail to apply them in a consistent way. These findings highlight the need to identify how and when LLMs' stated beliefs align with their simulated behavior, allowing researchers to use LLM-based agents appropriately in behavioral studies.
Citations
- Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy
- The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
- Can LLM Agents Maintain a Persona in Discourse?
- Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- LLM-based Human Simulations Have Not Yet Been Reliable
- User Simulation in the Era of Generative AI: User Modeling, Synthetic Data Generation, and System Evaluation
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Controllable Context Sensitivity and the Knob Behind It
- Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
- Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
- Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- On the limits of agency in agent-based models
- When Context Leads but Parametric Memory Follows in Large Language Models
- The Llama 3 Herd of Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
- Uncovering Name-Based Biases in Large Language Models Through Simulated Trust Game
- Characteristic AI Agents via Large Language Models
- Using large language models to generate silicon samples in consumer and marketing research: Challenges, opportunities, and guidelines
- The Illusion of Artificial Inclusion
- LARP: Language-Agent Role Play for Open-World Games
- InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
- Character-LLM: A Trainable Agent for Role-Playing
- EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Generative Agents: Interactive Simulacra of Human Behavior
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Out of One, Many: Using Language Models to Simulate Human Samples
- Trust, Reciprocity, and Social History
- OpenAI o1 System Card
Related