Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
2025/03/03 by Kanishk Gandhi, Gandhi, Kanishk, Ayush Chakravarthy +7 · 22 voices · 148 citations
#cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2503.01307
Abstract
Test-time inference has emerged as a powerful paradigm for enabling language models to ``think'' longer and more carefully about complex challenges, much like skilled human experts. While reinforcement learning (RL) can drive self-improvement in language models on verifiable tasks, some models exhibit substantial gains while others quickly plateau. For instance, we find that Qwen-2.5-3B far exceeds Llama-3.2-3B under identical RL training for the game of Countdown. This discrepancy raises a critical question: what intrinsic properties enable effective self-improvement? We introduce a framework to investigate this question by analyzing four key cognitive behaviors -- verification, backtracking, subgoal setting, and backward chaining -- that both expert human problem solvers and successful language models employ. Our study reveals that Qwen naturally exhibits these reasoning behaviors, whereas Llama initially lacks them. In systematic experimentation with controlled behavioral datasets, we find that priming Llama with examples containing these reasoning behaviors enables substantial improvements during RL, matching or exceeding Qwen's performance. Importantly, the presence of reasoning behaviors, rather than correctness of answers, proves to be the critical factor -- models primed with incorrect solutions containing proper reasoning patterns achieve comparable performance to those trained on correct solutions. Finally, leveraging continued pretraining with OpenWebMath data, filtered to amplify reasoning behaviors, enables the Llama model to match Qwen's self-improvement trajectory. Our findings establish a fundamental relationship between initial reasoning behaviors and the capacity for improvement, explaining why some language models effectively utilize additional computation while others plateau.
Citations
Cited by
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- Embarrassingly Simple Self-Distillation Improves Code Generation
- Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
- Base Models Know How to Reason, Thinking Models Learn When
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones
- Bootstrapping Task Spaces for Self-Improvement
- Cognitive models can reveal interpretable value trade-offs in language models
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Diversity or Precision? A Deep Dive into Next Token Prediction
- Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
- Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- Trust-Region Adaptive Policy Optimization
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- Wait, Wait, Wait... Why Do Reasoning Models Loop?
- Metacognitive Sensitivity for Test-Time Dynamic Model Selection
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
- Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
- Rectifying LLM Thought from Lens of Optimization
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- ScRPO: From Errors to Insights
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
- Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
- When Models Outthink Their Safety: Mitigating Self-Jailbreak in Large Reasoning Models with Chain-of-Guardrails
- The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
- Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
- IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
- Representation-Based Exploration for Language Models: From Test-Time to Post-Training
- More Than One Teacher: Adaptive Multi-Guidance Policy Optimization for Diverse Exploration
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
- Don't Just Fine-tune the Agent, Tune the Environment
- Skill-Targeted Adaptive Training
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
- Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
- Modeling Student Learning with 3.8 Million Program Traces
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- What Can You Do When You Have Zero Rewards During RL?
- MASH: Modeling Abstention via Selective Help-Seeking
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Scaling Generalist Data-Analytic Agents
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
- PIPer: On-Device Environment Setup via Online Reinforcement Learning
- Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
- Think Socially via Cognitive Reasoning
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- PSRT: Accelerating LRM-based Guard Models via Prefilled Safe Reasoning Traces
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- RLP: Reinforcement as a Pretraining Objective
- Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
- Proximal Supervised Fine-Tuning
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Can GRPO Boost Complex Multimodal Table Understanding?
- The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
- Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
- Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
- Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
- LegalΔ: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- SSRL: Self-Search Reinforcement Learning
- From Diagnosis to Improvement: Probing Spatio-Physical Reasoning in Vision Language Models
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- The SMeL Test: A simple benchmark for media literacy in language models
- Test-time Prompt Intervention
- Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities
- Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
- LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
- Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
- Agentic Reinforced Policy Optimization
- PurpCode: Reasoning for Safer Code Generation
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
- URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
- Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
- One Token to Fool LLM-as-a-Judge
- What Factors Affect LLMs and RLLMs in Financial Question Answering?
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Theoretical Modeling of LLM Self-Improvement Training Dynamics Through Solver-Verifier Gap
- Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
- AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
- Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Discussions
- Cognitive Behaviors That Enable Self-Improving Reasoners [hn, 279 points, 103 comments]
- it was this paper that was the lightbulb moment for me arxiv.org/abs/2503.013... [bsky, 7 points, 1 comments]
- PPO/GRPO/etc have no explicit mechanism to promote novelty or diversity, and there is evidence that current approaches may only be amplifying/reinforcing behavior already present in the base model (se [bsky, 5 points, 1 comments]
- Thanks for sharing, this looks like a really neat set of results! Feels potentially consistent with some of the findings from arxiv.org/abs/2503.01307 which I thought was a really enlightening paper t [bsky, 4 points, 1 comments]
- The birth of researchers studying the psychology of AIs is really fascinating to watch arxiv.org/abs/2503.01307 [bsky, 4 points, 0 comments]
- "Cognitive Behaviors That Enable Self-Improving Reasoners" Self-improving AI might help humans get better at thinking too. But we must be careful, as AIs might create confusing ways to understand the [bsky, 2 points, 0 comments]
- 13/13 Paper at arxiv.org/abs/2503.01307 [bsky, 2 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 [comments] [240 points] [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 (https://news.ycombinator.com/item?id=43275193) [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 (http://news.ycombinator.com/item?id=43275193) [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 (http://news.ycombinator.com/item?id=43275193) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2503.01307 自己改善型推論を可能にする認知行動について述べられた海外論文。 効果的なSTaR(Search, Test, and Revise)の4つの習慣が紹介されています。 論文では、認知的な側面から自己改善のメカニズムを分析しています。 [bsky, 0 points, 0 comments]
- Cognitive Behaviors that Enable Self-Improving Reasoners arxiv.org/abs/2503.01307 [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Cognitive Behaviors That Enable Self-Improving Reasoners [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 (https://news.ycombinator.com/item?id=43275193) [bsky, 0 points, 0 comments]
- 1) verification (systematic error-checking) 2) backtracking (abandoning failing approaches) 3) subgoal setting (decomposing problems) 4) backward chaining [bsky, 0 points, 0 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3ljopfypqrfj2 [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners #HackerNews https://arxiv.org/abs/2503.01307 [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 https://news.ycombinator.com/item?id=43275193 [bsky, 0 points, 0 comments]
- Cognitive Behaviors That Enable Self-Improving Reasoners https://arxiv.org/abs/2503.01307 [bsky, 0 points, 0 comments]
Related