Reasoning Models Don't Always Say What They Think
2025/05/08 by Yanda Chen, Chen, Yanda, Joe Benton +27 · 6 voices · 95 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2505.05410
openalex publication_date 2025/05/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.
Cited by
- Compared to What? Baselines and Metrics for Counterfactual Prompting
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Design Theater: A Benchmark for Generative UI
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- LIR3AG: A Lightweight Rerank Reasoning Strategy Framework for Retrieval-Augmented Generation
- Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- State over Tokens: Characterizing the Role of Reasoning Tokens
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
- Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
- On the Regulatory Potential of User Interfaces for AI Agent Governance
- Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment
- When to Think and When to Look: Uncertainty-Guided Lookback
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Investigating CoT Monitorability in Large Reasoning Models
- Revealing AI Reasoning Increases Trust but Crowds Out Unique Human Knowledge
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
- Reasoning Models Sometimes Output Illegible Chains of Thought
- Position: Evaluation Scores Are Perishable Knowledge Claims
- A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
- On the Use of Large Language Models for Qualitative Synthesis
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Optimizing Retrieval for RAG via Reinforced Contrastive Learning
- A Pragmatic Way to Measure Chain-of-Thought Monitorability
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Mapping Faithful Reasoning in Language Models
- Modeling Hierarchical Thinking in Large Reasoning Models
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
- CourtGuard: A Local, Multiagent Prompt Injection Classifier
- Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
- The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
- Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
- Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
- ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
- Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
- LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
- Output Supervision Can Obfuscate the Chain of Thought
- Superficial Beliefs in LLM Decision-Making
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Disentangling Recall and Reasoning in Transformer Models through Layer-wise Attention and Activation Analysis
- Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
- Large Reasoning Models Learn Better Alignment from Flawed Thinking
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
- Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
- From Faithfulness to Correctness: Generative Reward Models that Think Critically
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- Tracing Uncertainty in Language Model "Reasoning"
- Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- OpenAI's GPT-OSS-20B Model and Safety Alignment Issues in a Low-Resource Language
- Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
- Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
- Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
- LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Code Like Humans: A Multi-Agent Solution for Medical Coding
- MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine
- Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
- AI reasoning effort predicts human decision time in content moderation
- Reliable Weak-to-Strong Monitoring of LLM Agents
- The AI in the Mirror: LLM Self-Recognition in an Iterated Public Goods Game
- Lexical Hints of Accuracy in LLM Reasoning Chains
- LLM-Assisted Functional Gene Annotation
- Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
- The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
Discussions
- Now, Anthropic, which has a very strong interest in proving chain of thought works, has released a paper saying that it doesn't. arxiv.org/abs/2505.054... [bsky, 8 points, 1 comments]
- Reasoning Models Don't Always Say What They Think arxiv.org/abs/2505.05410 多肢選択問題に、解答についてのヒントを付与した問題とヒントを付与していない問題をLLMに解かせてみて、解答が変化した問題に注目します。この時、推論過程ではヒントを参照しているはずと考えられますが、推論過程を出力しろと要求しても「ヒントを利用した」とは [bsky, 3 points, 1 comments]
- Reasoning Models Don't Always Say What They Think [hn, 2 points, 0 comments]
- Oops! #FaithfullyWaitingForAGI 😜 arxiv.org/abs/2505.05410 [bsky, 1 points, 0 comments]
- Reasoning Models Don't Always Say What They Think [hn, 1 points, 0 comments]
- Same topic, a little bit early reference, and from the Anthropic team. arxiv.org/abs/2505.05410 [bsky, 0 points, 0 comments]
Related