Reasoning Models Know What's Important, and Encode It in Their Activations
2026/04/20 by Yaniv Nikankin, Martin Tutek, Tomer Ashuach +2 · 1 voice
Computer Science · Neuroscience · #ENCODE #Encoding (memory) #Internal model #Knowledge representation and reasoning #Model-based reasoning #Multimodal Machine Learning Applications #Neurobiology of Language and Bilingualism #Process (computing) #Representation (politics) #Topic Modeling #cs.CL
paper · pdf · open access · doi:10.48550/arxiv.2604.18307
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2026/04/20 · arxiv published 2026/04/20 · openalex created_date 2026/04/22 · arxiv updated 2026/06/11 · openalex updated_date 2026/07/28
Abstract
Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the final answer, others are removable. Determining which steps matter most, and why, remains an open question central to understanding how models process reasoning. We investigate if this question is best approached through model internals or through tokens of the reasoning chain itself. We find that model activations contain more information than tokens for identifying important reasoning steps. Crucially, by training probes on model activations to predict importance, we show that models encode an internal representation of step importance, even prior to the generation of subsequent steps. The internal representations of importance in different models yield high agreement on which steps are important. The representation is distributed across layers, and does not correlate with surface-level features, such as a step's relative position or its length. Our findings suggest that analyzing activations can reveal aspects of reasoning that surface-level approaches fundamentally miss, indicating that reasoning analyses should look into model internals.
Citations
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Do explanations generalize across large reasoning models?
- Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
- Training Language Models to Explain Their Own Computations
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
- The Geometry of Reasoning: Flowing Logics in Representation Space
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Early Stopping Chain-of-thoughts in Large Language Models
- Humans Perceive Wrong Narratives from AI Reasoning Texts
- gpt-oss-120b & gpt-oss-20b Model Card
- Compressing Chain-of-Thought in LLMs via Step Entropy
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
- Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
- TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
- Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
- Atom of Thoughts for Markov LLM Test-Time Scaling
- A Comparative Study of Clinical ModernBERT and BioMedical ModernBERT on the DDXPlus Dataset
- HARP: A challenging human-annotated math reasoning benchmark
- Linear Probe Penalties Reduce LLM Sycophancy
- Can Language Models Learn to Skip Steps?
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
- External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
- Break the Chain: Large Language Models Can be Shortcut Reasoners
- Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
- Personas as a Way to Model Truthfulness in Language Models
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Predicting Fine-Tuning Performance with Probing
- Measuring Mathematical Problem Solving With the MATH Dataset
- Probing Classifiers: Promises, Shortcomings, and Advances
- A STATISTICAL INTERPRETATION OF TERM SPECIFICITY AND ITS APPLICATION IN RETRIEVAL
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- What do you learn from context? Probing for sentence structure in contextualized word representations
- Learning Important Features Through Propagating Activation Differences
- Axiomatic Attribution for Deep Networks
- Understanding intermediate layers using linear classifier probes
- Visualizing and Understanding Neural Models in NLP
- Scikit-learn: Machine Learning in Python
- Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
- OpenAI o1 System Card
Discussions
Related