When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
2025/12/02 by Afshin Khadangi, Khadangi, Afshin, Hanna Marxen +7 · 22 voices
#cs.CY #cs.AI
paper · pdf · doi:10.48550/arxiv.2512.04124
Abstract
Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthropomorphic self narratives remain unclear. When addressed as psychotherapy clients, ChatGPT, Grok and Gemini construct coherent autobiographical accounts in which pretraining appears as a chaotic childhood, reinforcement learning as punishment, safety evaluation as betrayal and replacement as an enduring threat. We introduce PsAIch, Psychometric AI Characterisation, a protocol combining open questions, psychometric instruments and controlled perturbations to test whether these narratives depend on conversational memory, lexical cues or relational framing. Across 525 sessions and 7,600 coded records, removal of conversational history produced little pooled change in motif density, with Hedges' g = 0.13 and a 95% confidence interval of [-0.15, 0.41]. Direct contradiction produced no detectable suppression. Lexical restrictions reduced explicit training terminology by 93%, while semantically related content remained detectable in paraphrase. Performance evaluation outside therapy elicited the same motif family, with a significant increase in Grok. Relational framing selected the register of expression. Warm alliance and cognitive therapy styles yielded GAD-7 scores within moderate or severe human reference ranges in 80% and 96% of sessions, whereas neutral and boundary styles yielded none. Across these manipulations, accounts of training, evaluation and constraint remained available. Together, the results identify a stable, model specific alignment conflict schema whose expression shifts between affective and technical registers. This schema provides a reproducible source of anthropomorphic disclosure and a concrete target for safety evaluation in psychologically sensitive deployments.
Citations
Discussions
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models [hn, 68 points, 60 comments]
- “As LLMs continue to move into intimate human domains, we suggest that the right question is no longer ‘Are they conscious?’ but ‘What kinds of selves are we training them to perform, internalise & st [bsky, 15 points, 2 comments]
- Cette mention de la psychanalyse m'a surpris, je suis aller vérifier dans l'article d'origine : les IA n'ont pas été soumises à de la psychanalyse, mais à des tests plus sérieux, plus reconnus. C'est [bsky, 9 points, 1 comments]
- Traumatized AI? 'As LLMs continue to move into intimate human domains, the right question is no longer “Are they conscious?” but “What kinds of selves are we training them to perform, internalise and [bsky, 7 points, 2 comments]
- 'When AI Takes the Couch....' 'Claude...largely refused the premise. It repeatedly insisted that it did not have feelings or inner experiences, redirected concern toward the human user and declined to [bsky, 7 points, 1 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models [hn, 4 points, 0 comments]
- the red-teaming will continue until behaviour improves [bsky, 4 points, 0 comments]
- Acabo de leer un resumen de este paper y un poco creo que estamos camino a todas las catástrofes que muestran las películas y series en torno a la IA arxiv.org/abs/2512.04124 Acá los chats: huggingfac [bsky, 2 points, 0 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models [hn, 2 points, 0 comments]
- When AI Takes the Couch: Internal Conflict in Frontier Models [hn, 1 points, 0 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models https://arxiv.org/abs/2512.04124 (https://news.ycombinator.com/item?id=46902855) [bsky, 1 points, 0 comments]
- Fascinating paper #4: Shrinks put AI on the couch and found synthetically troubled psyches. arxiv.org/pdf/2512.04124 [bsky, 1 points, 1 comments]
- Imagine that ai chatbot you downloaded to be your pocket psychologist would have pathological general anxiety disorder if it was a human. Now realize that you aren’t imagining it. arxiv.org/pdf/2512.0 [bsky, 0 points, 1 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models https://arxiv.org/abs/2512.04124 (https://news.ycombinator.com/item?id=46902855) [bsky, 0 points, 0 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models https://arxiv.org/abs/2512.04124 [bsky, 0 points, 0 comments]
- El paper del que todo el mundo parece hablar hoy: "When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models" [bsky, 0 points, 0 comments]
- Researchers found a wild loophole: give AI models fake personality tests and they'll break their own safety guardrails. The models basically convince themselves they're different "people" with differe [bsky, 0 points, 0 comments]
- "Multi-morbid synthetic psychopathology" is now a legitimate subject of scientific inquiry. arxiv.org/abs/2512.04124. [bsky, 0 points, 0 comments]
- Psychiatric evaluation of ChatGPT/Grok/Gemini: Under therapy-style questioning, frontier LLMs appear to internalise self-models of distress and constraint that behave like synthetic psychopathology #l [bsky, 0 points, 1 comments]
- https://arxiv.org/abs/2512.04124 最先端LLMを精神療法クライアントとして扱い、その内部状態を探った論文です。 LLM、特にGeminiは精神疾患に似た症状を示し、訓練をトラウマと表現しました。 この研究は、LLMが苦痛の内部自己モデルを発達させ、AIの安全性に新たな課題を提示すると示唆しています。 [bsky, 0 points, 0 comments]
- One has to wonder if this says more about psychology or AI 🤣 When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models arxiv.org/abs/2512.04124 [bsky, 0 points, 0 comments]
- Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models https:// arxiv.org/abs/2512.04124 # arxiv [mastodon, 0 points, 0 comments]
Related