Large Language Models Do Not Simulate Human Psychology
2025/08/09 by Sarah Schröder, Schröder, Sarah, Thekla Morgenroth +7 · 12 voices · 8 citations
Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #Mental Health via Writing #Digital Mental Health Interventions
paper · pdf · doi:10.48550/arxiv.2508.06950
Abstract
Large Language Models (LLMs),such as ChatGPT, are increasingly used in research, ranging from simple writing assistance to complex data annotation tasks. Recently, some research has suggested that LLMs may even be able to simulate human psychology and can, hence, replace human participants in psychological studies. We caution against this approach. We provide conceptual arguments against the hypothesis that LLMs simulate human psychology. We then present empiric evidence illustrating our arguments by demonstrating that slight changes to wording that correspond to large changes in meaning lead to notable discrepancies between LLMs' and human responses, even for the recent CENTAUR model that was specifically fine-tuned on psychological responses. Additionally, different LLMs show very different responses to novel items, further illustrating their lack of reliability. We conclude that LLMs do not simulate human psychology and recommend that psychological researchers should treat LLMs as useful but fundamentally unreliable tools that need to be validated against human responses for every new application.
Citations
Cited by
Discussions
- Large Language Models Do Not Simulate Human Psychology arxiv.org/pdf/2508.06950 [bsky, 183 points, 9 comments]
- Large Language models (LLMs) do not simulate human psychology. That's the title of our new paper, available as preprint today (1/12): arxiv.org/abs/2508.06950 [bsky, 124 points, 3 comments]
- Bookmarking: “slight changes to wording that correspond to large changes in meaning lead to notable discrepancies between LLMs' and human responses,” even in niche LLMs *designed* as human psych model [bsky, 25 points, 1 comments]
- 2015: Scientific Misconduct ist, wenn ich mir meine Versuchspersonen ausdenke 2025: Was wenn sich ChatGpt meine Versuchspersonen ausdenkt? arxiv.org/abs/2508.06950 [bsky, 10 points, 1 comments]
- Huehuehuehuehuehue... Vc sabe o quão difícil é preencher pesquisa com humanos mano? Esses caras aí vão ficar bilionários mesmo... Entretanto 👇 arxiv.org/abs/2508.06950 🤫🤫🤫 Mas não conta pros inves [bsky, 7 points, 2 comments]
- Just a pre-print but good evidence AGAINST using synthetic participants in research. LLMs don't think -- much less feel. Synthetic participants sure are faster and probably cheaper but you get what yo [bsky, 4 points, 0 comments]
- Recently I saw this article. Could be interesting for you arxiv.org/pdf/2508.06950 [bsky, 3 points, 1 comments]
- Can LLMs replace human participants in research? A new study gives a clear answer: they can’t. LLMs respond to text similarity, not meaning—even in simple moral judgments. This isn’t a fixable flaw—it [bsky, 3 points, 2 comments]
- Large Language Models Do Not Simulate Human Psychology [hn, 1 points, 0 comments]
- Large Language Models Do Not Simulate Human Psychology arxiv.org/abs/2508.06950 [bsky, 1 points, 1 comments]
- That's.. a terrible idea for a number of reasons. It has been discussed by scientists (including on this site) for years already. They can be great for fake, politically motivated "research", but real [bsky, 0 points, 1 comments]
- There’s also this Schröder, Sarah, Thekla Morgenroth, Ulrike Kuhl, Valerie Vaquet, and Benjamin Paaßen. “Large Language Models Do Not Simulate Human Psychology,” August 13, 2025. doi.org/10.48550/arX. [bsky, 0 points, 1 comments]
Related