Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
2025/04/07 by Anqi Zhang, Zhang, Anqi, Yulin Chen +13 · 1 voice · 73 citations
Computer Science · #Advanced Graph Neural Networks #Correctness #Deductive reasoning #ENCODE #Exploit #Inference #Intelligent Tutoring Systems and Adaptive Learning #Logical consequence #Model-based reasoning #Reasoning system #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2504.05419
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/04/07 · arxiv published 2025/04/07 · arxiv updated 2025/04/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Reasoning models have achieved remarkable performance on tasks like math and logical reasoning thanks to their ability to search during reasoning. However, they still suffer from overthinking, often performing unnecessary reasoning steps even after reaching the correct answer. This raises the question: can models evaluate the correctness of their intermediate answers during reasoning? In this work, we study whether reasoning models encode information about answer correctness through probing the model's hidden states. The resulting probe can verify intermediate answers with high accuracy and produces highly calibrated scores. Additionally, we find models' hidden states encode correctness of future answers, enabling early prediction of the correctness before the intermediate answer is fully formulated. We then use the probe as a verifier to decide whether to exit reasoning at intermediate answers during inference, reducing the number of inference tokens by 24% without compromising performance. These findings confirm that reasoning models do encode a notion of correctness yet fail to exploit it, revealing substantial untapped potential to enhance their efficiency.
Citations
Cited by
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Entropy Sentinel: Probing Entropy Traces for LLM Monitoring
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
- Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
- Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
- Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
- LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
- ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- Therefore I am. I Think
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- Internal states before wait modulate reasoning patterns
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- Entropy After ⟨
/Think ⟩ for reasoning model early exiting - Intra-request branch orchestration for efficient LLM reasoning
- Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
- SCI-Verifier: Scientific Verifier with Thinking
- SpecExit: Accelerating Large Reasoning Model via Speculative Exit
- Learning to Ponder: Adaptive Reasoning in Latent Space
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- Tracing Uncertainty in Language Model "Reasoning"
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- Learning to Reason Across Parallel Samples for LLM Reasoning
- Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- The Geometries of Truth Are Orthogonal Across Tasks
- BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
- CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
- Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning Models
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- Real-Time Progress Prediction in Reasoning Language Models
- Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement
- The Price of a Second Thought: On the Evaluation of Reasoning Efficiency in Large Language Models
- Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
- Thought calibration: Efficient and confident test-time scaling
- The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
- Let LRMs Break Free from Overthinking via Self-Braking Tuning
- FlashThink: An Early Exit Method For Efficient Reasoning
- Reasoning Models Better Express Their Confidence
- Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
- Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation
- Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
- Conformal Thinking: Risk Control for Reasoning on a Compute Budget
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
- TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
- The Value Axis: Language Models Encode Whether They're on the Right Track
- Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
- Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
- LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
- Reasoning Models Know What's Important, and Encode It in Their Activations
- Dynamic Early Exit in Reasoning Models
- The Geometry of Self-Verification in a Task-Specific Reasoning Model
Discussions
Related