vix.ing · top · new · best · stats · spec

Persuading large language models to comply with objectionable requests

2026/05/19 by Lennart Meincke, Dan Shapiro, Angela L. Duckworth +4 · 1 voice
Neuroscience · Psychology · Social Sciences · #Deception detection and forensic psychology #Psychology of Moral and Emotional Judgment #Psychology of Social Influence

paper · doi:10.1073/pnas.2535868123

openalex publication_date 2026/05/19 · openalex created_date 2026/05/20 · openalex updated_date 2026/07/31

Abstract

Are large language models (LLMs) susceptible to the same persuasive appeals as humans? We tested whether classic persuasion principles (authority, commitment, liking, reciprocity, scarcity, social proof, and unity) could induce three widely used LLMs (GPT-5 mini, Claude Haiku 4.5, and Gemini 3 Flash) to comply with requests to assist with the synthesis of regulated substances. Across 126,000 conversations, persuasion principles increased compliance from 35.3% (at baseline) to 51.3% (using any principle). Although LLMs are not human, these findings underscore their parahuman (i.e., humanlike) nature and reveal the risk of manipulation by malicious users seeking to circumvent safety guardrails.

Citations

Discussions

Related