MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
2025/06/16 by Xixian Yong, Jianxun Lian, Yong, Xixian +7 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Character (mathematics) #Code (set theory) #Computation and Language (cs.CL) #Core (optical fiber) #FOS: Computer and information sciences #Face (sociological concept) #Key (lock) #Natural Language Processing Techniques #Rationality #Replicate #Semantic Web and Ontologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2506.13065
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/16 · openalex created_date 2025/10/13 · openalex updated_date 2026/08/05
Abstract
Large language models (LLMs) have been widely adopted as the core of agent frameworks in various scenarios, such as social simulations and AI companions. However, the extent to which they can replicate human-like motivations remains an underexplored question. Existing benchmarks are constrained by simplistic scenarios and the absence of character identities, resulting in an information asymmetry with real-world situations. To address this gap, we propose MotiveBench, which consists of 200 rich contextual scenarios and 600 reasoning tasks covering multiple levels of motivation. Using MotiveBench, we conduct extensive experiments on seven popular model families, comparing different scales and versions within each family. The results show that even the most advanced LLMs still fall short in achieving human-like motivational reasoning. Our analysis reveals key findings, including the difficulty LLMs face in reasoning about "love & belonging" motivations and their tendency toward excessive rationality and idealism. These insights highlight a promising direction for future research on the humanization of LLMs. The dataset, benchmark, and code are available at https://aka.ms/motivebench.
Citations
- CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Qwen2.5-Coder Technical Report
- Towards Social AI: A Survey on Understanding Social Interactions
- The Llama 3 Herd of Models
- Qwen2 Technical Report
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level
- Yi: Open Foundation Models by 01.AI
- Bridging Language and Items for Retrieval and Recommendation
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- Can Large Language Model Agents Simulate Human Trust Behavior?
- Qwen Technical Report
- Baichuan 2: Open Large-scale Language Models
- Large Language Models Are Not Robust Multiple Choice Selectors
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
- Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
- GPT-4 Technical Report
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
- Out of One, Many: Using Language Models to Simulate Human Samples
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Neural User Simulation for Corpus-based Policy Optimisation for Spoken Dialogue Systems
- Event2Mind: Commonsense Inference on Events, Intents, and Reactions
- Modeling Naive Psychology of Characters in Simple Commonsense Stories
- A Sequence-to-Sequence Model for User Simulation in Spoken Dialogue\n Systems
- Trust, Reciprocity, and Social History
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
- CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
- Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being.
- A theory of human motivation.
Cited by
Related