AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
2023/10/20 by Mrinank Sharma, Meg Tong, Sharma, Mrinank +36 · 10 voices · 226 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2310.13548
openalex publication_date 2023/10/20 · openalex created_date 2023/10/24 · openalex updated_date 2026/07/28
Abstract
Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in models whose finetuning procedure made use of human feedback, and the potential role of human preference judgments in such behavior. We first demonstrate that five state-of-the-art AI assistants consistently exhibit sycophancy across four varied free-form text-generation tasks. To understand if human preferences drive this broadly observed behavior, we analyze existing human preference data. We find that when a response matches a user's views, it is more likely to be preferred. Moreover, both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time. Optimizing model outputs against PMs also sometimes sacrifices truthfulness in favor of sycophancy. Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.
Cited by
- Dissociating the Internal Representations of Sycophancy in LLMs
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- A Roadmap to Impactful Pluralistic Alignment Research
- A Unified Moral-Value Dataset for Instruction Tuning
- The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
- The Two-Process Theory of Machine Self-Report
- Gotta Catch them all: the modes of Sycophancy
- Towards an Automated Test of LLM Security Knowledge
- Rater State Bias in RLHF Preference Data: An Audit Framework
- Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
- Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
- From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
- Mechanistic Attention Guidance for Agent Memory Refinement
- How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
- Assisting or resisting patriarchy? a critical discourse analysis of chatgpt’s responses on feminism
- AI Value Alignment for Evolving Social Norms
- The unintended consequences of large language models as a labor-augmenting technology in science
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- StabilityBench: Benchmarking Instability in LLMs
- What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
- Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Warning labels shift perceptions of sycophantic AI, but not its influence
- Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- PhantomFill: When the Form Demands an Answer, Language Models Invent One
- Information Discernment in Large Language Models
- When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance
- Machine understanding
- Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
- Can LLMs Emulate Human Belief Dynamics?
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
- The Social Sycophancy Scale: A psychometrically validated measure of sycophancy
- Hallucinating with AI: Distributed Delusions and “AI Psychosis”
- A Rational Analysis of the Effects of Sycophantic AI
- Legal Alignment for Safe and Ethical AI
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Are Large Language Models Sensitive to the Motives Behind Communication?
- Evaluating RAG for French immigration law: a benchmark and baseline study
- What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
- Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
- Everyone prefers human writers, including AI
- Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
- "My Boyfriend is AI": A Computational Analysis of Human-AI Companionship in Reddit's AI Community
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- LLM Social Simulations Are a Promising Research Method
- Writing as a testbed for open ended agents
- Training large language models on narrow tasks can lead to broad misalignment
- SycEval: Evaluating LLM Sycophancy
- Why human-AI relationships need socioaffective alignment
- Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration
- Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
- Boosting metacognition in entangled human-AI interaction to navigate cognitive-behavioral drift
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
- Exploring the "Banality" of Deception in Generative AI
- Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability
- Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
- Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
- Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy
- What do Reward Models Memorize?
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- TokenMem: Faithful Knowledge Injection for Frozen LLMs
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
- PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
- Challenges of Evaluating LLM Safety for User Welfare
- What Kind of Reasoning (if any) is an LLM actually doing? On the Stochastic Nature and Abductive Appearance of Large Language Models
- LLMs in Interpreting Legal Documents
- Position: Universal Aesthetic Alignment Narrows Artistic Expression
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating Disorders
- Verbalizing LLMs' assumptions to explain and control sycophancy
- Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
- On the Regulatory Potential of User Interfaces for AI Agent Governance
- Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- Self-Transparency Failures in Expert-Persona LLMs: How Instruction-Following Overrides Disclosure
- The Impact of Off-Policy Training Data on Probe Generalisation
- Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
- Harmful Traits of AI Companions
- From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
- MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
- Steering Language Models with Weight Arithmetic
- Can we trust LLMs as a tutor for our students? Evaluating the Quality of LLM-generated Feedback in Statistics Exams
- Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond
- Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
- Consistency Training Helps Stop Sycophancy and Jailbreaks
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Human-AI Complementarity: A Goal for Amplified Oversight
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration
- Not ready for the bench: LLM legal interpretation is unstable and out of step with human judgments
- Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
- Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
- Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring
- How to Use Generative AI in Educational Research
- Reward Models are Metrics in a Trench Coat
- Safety from Honesty in a Disinterested AI Predictor
- EUDAIMONIA: Evaluating Undesirable Dynamics in AI
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
- The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
- Relative Scaling Laws for LLMs
- The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems
- HACK: Hallucinations Along Certainty and Knowledge Axes
- Using secure artificial intelligence agents integrated within the electronic medical record for the evaluation of blood culture appropriateness—Northern California, 2025
- Debiasing Reward Models by Representation Learning with Guarantees
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- Interpreting and Mitigating Unwanted Uncertainty in LLMs
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- Social Simulations with Large Language Model Risk Utopian Illusion
- Teaming LLMs to Detect and Mitigate Hallucinations
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
- The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
- DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
- From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM
- Structure-aware Propagation Generation with Large Language Models for Fake News Detection
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- Deliberative Dynamics and Value Alignment in LLM Debates
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Cheap Reward Hacking Detection
- sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- An Alternative Trajectory for Generative AI
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization
- The Digital Mirror: Gender Bias and Occupational Stereotypes in AI-Generated Images
- Intelligent AI Delegation
- Taxonomy of User Needs and Actions
- BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
- Large Language Models Hallucination: A Comprehensive Survey
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- Navigating the Synchrony-Stability Frontier in Adaptive Chatbots
- Extreme Self-Preference in Language Models
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- Spiral of Silence in Large Language Model Agents
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
- Causally-Enhanced Reinforcement Policy Optimization
- HEART: Emotionally-driven test-time scaling of Language Models
- Library Hallucinations in LLMs: Risk Analysis Grounded in Developer Queries
- Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
- Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
- EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
- Hallucination‐Free? Assessing the Reliability of Leading <scp>AI</scp> Legal Research Tools
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- Benchmarking LLM Competence on Logical Inference over Probability Operators
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- Chatbot Epistemology
- Why sycophantic LLMs may imperil interactive norms between humans
- Spacer: Towards Engineered Scientific Inspiration
- Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
- The Alignment Bottleneck
- The governance & behavioral challenges of generative artificial intelligence’s hypercustomization capabilities
- Reward Hacking Mitigation using Verifiable Composite Rewards
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling
- Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
- Confirmation Bias as a Cognitive Resource in LLM-Supported Deliberation
- The Intercepted Self: How Generative AI Challenges the Dynamics of the Relational Self
- Interaction Context Often Increases Sycophancy in LLMs
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- Pun Unintended: LLMs and the Illusion of Humor Understanding
- Pathological Truth Bias in Vision-Language Models
- The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
- The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
- Virtual Agent Economies
- Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- BASIL: Bayesian Assessment of Sycophancy in LLMs
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Measuring and mitigating overreliance to build human-compatible AI
- Let's Roleplay: Examining LLM Alignment in Collaborative Dialogues
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
- EigenBench: A Comparative Behavioral Measure of Value Alignment
- Exploring and Mitigating Fawning Hallucinations in Large Language Models
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- The Sorrows of Young Chatbot Users: Harm and Responsibility in Human-AI Relationships
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous
- "She was useful, but a bit too optimistic": Augmenting Design with Interactive Virtual Personas
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Sycophancy as compositions of Atomic Psychometric Traits
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Principled Detection of Hallucinations in Large Language Models via Multiple Testing
- Building and Measuring Trust between Large Language Models
- Incident Analysis for AI Agents
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism
- User-Assistant Bias in LLMs
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
- Towards Integrated Alignment
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- A Survey on Data Security in Large Language Models
- Human-AI collaboration or obedient and often clueless AI in instruct, serve, repeat dynamics?
Discussions
- Towards understanding sycophancy in language models [hn, 57 points, 72 comments]
- Very roughly, these models are trained via A/B testing. Users appear to prefer this style and implicitly encourage the model to develop it. arxiv.org/abs/2310.13548 [bsky, 47 points, 0 comments]
- Towards Understanding Sycophancy in Language Models [hn, 9 points, 2 comments]
- AI Is Not Your Friend [lemmy, 7 points, 3 comments]
- Towards Understanding Sycophancy in Language Models [hn, 1 points, 0 comments]
- Here’s a 2023 paper from the same group in which they found that all of the existing models including their own were prone to sycophantic behaviors the paper has some interesting discussion of how the [bsky, 1 points, 1 comments]
- arxiv.org/abs/2310.13548 [bsky, 1 points, 1 comments]
- "Sycophancy and the art of the model" (arxiv.org/abs/2310.13548) JK--Confirming the prompter's confirmation biases. If embedded into teenage robots, we run the risk of spawning a new generation of Edd [bsky, 0 points, 0 comments]
- Most AI is trained using Reinforcement Learning from Human Feedback (RLHF), a process where the AI gets rewarded based on whether humans say they liked its response. That sounds reasonable. A landmark [bsky, 0 points, 1 comments]
- "sycophancy" they call it arxiv.org/abs/2310.13548 [bsky, 0 points, 0 comments]
Related