From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
2025/06/20 by Zhicheng Lin, Lin, Zhicheng · 3 citations
Neuroscience · Psychology · Social Sciences · #Neurobiology of Language and Bilingualism #Mental Health via Writing #Computational and Text Analysis Methods
paper · pdf · doi:10.48550/arxiv.2506.16697
Abstract
Large language models (LLMs) are rapidly being adopted across psychology, serving as research tools, experimental subjects, human simulators, and computational models of cognition. However, the application of human measurement tools to these systems can produce contradictory results, raising concerns that many findings are measurement phantoms--statistical artifacts rather than genuine psychological phenomena. In this Perspective, we argue that building a robust science of AI psychology requires integrating two of our field's foundational pillars: the principles of reliable measurement and the standards for sound causal inference. We present a dual-validity framework to guide this integration, which clarifies how the evidence needed to support a claim scales with its scientific ambition. Using an LLM to classify text may require only basic accuracy checks, whereas claiming it can simulate anxiety demands a far more rigorous validation process. Current practice systematically fails to meet these requirements, often treating statistical pattern matching as evidence of psychological phenomena. The same model output--endorsing "I am anxious"--requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model psychological constructs. Moving forward requires developing computational analogues of psychological constructs and establishing clear, scalable standards of evidence rather than the uncritical application of human measurement tools.
Citations
- Large Language Models as Psychological Simulators: A Methodological Guide
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Does It Make Sense to Speak of Introspection in Large Language Models?
- Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement
- LLMs Get Lost In Multi-Turn Conversation
- Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
- Towards LLMs Robustness to Changes in Prompt Format Styles
- Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
- R.U.Psycho? Robust Unified Psychometric Testing of Language Models
- On Benchmarking Human-Like Intelligence in Machines
- How to evaluate the cognitive abilities of LLMs
- Using natural language processing to analyse text data in behavioural science
- Position: Theory of Mind Benchmarks are Broken for Large Language Models
- Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
- Sense and Sensitivity: Evaluating the simulation of social dynamics via Large Language Models
- Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
- Does Prompt Formatting Have Any Impact on LLM Performance?
- Large-scale moral machine experiment on large language models
- One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity
- Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
- Larger and more instructable language models become less reliable
- Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
- Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges
- Evaluating Large Language Models with Psychometrics
- Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
- GPT-ology, Computational Models, Silicon Sampling: How should we think about LLMs in Cognitive Science?
- Synthetic Replacements for Human Survey Data? The Perils of Large Language Models
- Large Language Models Show Human-like Social Desirability Biases in Survey Responses
- A Philosophical Introduction to Language Models - Part II: The Way Forward
- Large language models that replace human participants can harmfully misportray and flatten identity groups
- AI for social science and social science of AI: A Survey
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- Challenging the Validity of Personality Tests for Large Language Models
- Do LLMs exhibit human-like response biases? A case study in survey design
- Techniques for supercharging academic writing with generative AI
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- Self-Assessment Tests are Unreliable Measures of LLM Personality
- PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
- Can Large Language Models Transform Computational Social Science?
- Can Large Language Models Transform Computational Social Science?
- Generative Agents: Interactive Simulacra of Human Behavior
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Evaluating large language models in theory of mind tasks
- Emotional Intelligence of Large Language Models
- Emergent Analogical Reasoning in Large Language Models
- Who is GPT-3? An Exploration of Personality, Values and Demographics
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- On Logical Inference over Brains, Behaviour, and Artificial Neural Networks
- Avoiding common machine learning pitfalls
- A Survey on Bias and Fairness in Machine Learning
- Objective Tests as Instruments of Psychological Theory
- The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
- Are Emergent Abilities of Large Language Models a Mirage?
- The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power.
- The Concept of Validity.
Cited by
Related