Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
2025/08/04 by Linan Yue, Yichao Du, Yue, Linan +19 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2508.02120
openalex publication_date 2025/08/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recently, Large Reasoning Models (LRMs) have gradually become a research hotspot due to their outstanding performance in handling complex tasks. Among them, DeepSeek R1 has garnered significant attention for its exceptional performance and open-source nature, driving advancements in the research of R1-style LRMs. Unlike traditional Large Language Models (LLMs), these models enhance logical deduction and decision-making capabilities during reasoning by incorporating mechanisms such as long chain-of-thought and self-reflection through reinforcement learning. However, with the widespread application of these models, the problem of overthinking has gradually emerged. Specifically, when generating answers, these models often construct excessively long reasoning chains with redundant or repetitive steps, which leads to reduced reasoning efficiency and may affect the accuracy of the final answer. To this end, various efficient reasoning methods have been proposed, aiming to reduce the length of reasoning paths without compromising model performance and reasoning capability. By reviewing the current research advancements in the field of efficient reasoning methods systematically, we categorize existing works into two main directions based on the lens of single-model optimization versus model collaboration: (1) Efficient Reasoning with Single Model, which focuses on improving the reasoning efficiency of individual models; and (2) Efficient Reasoning with Model Collaboration, which explores optimizing reasoning paths through collaboration among multiple models. Besides, we maintain a public GitHub repository that tracks the latest progress in efficient reasoning methods.
Citations
- Activation Steering for Chain-of-Thought Compression
- Controlling Thinking Speed in Reasoning Models
- AALC: Large Language Model Efficient Reasoning via Adaptive Accuracy-Length Control
- Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
- Optimizing Length Compression in Large Reasoning Models
- TagRouter: Learning Route to LLMs through Tags for Open-Domain Text Generation Tasks
- Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
- Efficient LLM Collaboration via Planning
- ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
- Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- On Reasoning Strength Planning in Large Reasoning Models
- What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
- How Far Are We from Optimal Reasoning Efficiency?
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
- Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- Long or short CoT? Investigating Instance-level Switch of Large Reasoning Models
- Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
- POSS: Position Specialist Generates Better Draft for Speculative Decoding
- Adaptive Graph Pruning for Multi-Agent Communication
- OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
- Answer Convergence as a Signal for Early Stopping in Reasoning
- The Price of a Second Thought: On the Evaluation of Reasoning Efficiency in Large Language Models
- Mitigating Overthinking in Large Reasoning Models via Manifold Steering
- Self-Route: Automatic Mode Switching via Capability Estimation for Efficient Reasoning
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
- Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition
- Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- VeriThinker: Learning to Verify Makes Reasoning Model Efficient
- Not All Tokens Are What You Need In Thinking
- Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
- Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
- TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
- Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
- Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
- ThinkSwitcher: When to Think Hard, When to Think Fast
- Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
- FlashThink: An Early Exit Method For Efficient Reasoning
- Think Only When You Need with Large Hybrid-Reasoning Models
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization
- Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
- Thinkless: LLM Learns When to Think
- AdaptThink: Reasoning Models Can Learn When to Think
- Efficient RL Training for Reasoning Models via Length-Aware Optimization
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
- Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
- Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
- Scalable Chain of Thoughts via Elastic Reasoning
- Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
- Accelerating Large Language Model Reasoning via Speculative Search
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
- Efficient Reasoning for LLMs through Speculative Chain-of-Thought
- Safety in Large Reasoning Models: A Survey
- Dynamic Early Exit in Reasoning Models
- Thought Manipulation: External Thought Can Be Efficient for Large Reasoning Models
- ToolRL: Reward is All Tool Learning Needs
- Guiding Reasoning in Small Language Models with LLM Assistance
- Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
- SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
- SEAL: Steerable Reasoning Calibration of Large Language Models for Free
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Think When You Need: Self-Adaptive Chain-of-Thought Learning
- Hawkeye:Efficient Reasoning with Model Collaboration
- TwT: Thinking without Tokens by Habitual Reasoning Distillation with Multi-Teachers' Guidance
- Efficient Inference for Large Reasoning Models: A Survey
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond
- Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- Unlocking General Long Chain-of-Thought Reasoning Capabilities of Large Language Models via Representation Engineering
- Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
- DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
- How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
- Chain of Draft: Thinking Faster by Writing Less
- Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
- H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
- To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
- LoRE-Merging: Exploring Low-Rank Estimation For Large Language Model Merging
- CoT-Valve: Length-Compressible Chain-of-Thought Tuning
- The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
- Reward-Guided Speculative Decoding for Efficient LLM Reasoning
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
- GraphRouter: A Graph-based Router for LLM Selections
- Uncovering Latent Chain of Thought Vectors in Language Models
- Learning Harmonized Representations for Speculative Sampling
- PEARL: Parallel Speculative Decoding with Adaptive Draft Length
- Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
- Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
- Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding
- External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
- Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- RouterBench: A Benchmark for Multi-LLM Routing System
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
- Theoretical guarantees on the best-of-n alignment policy
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
- Representation Engineering: A Top-Down Approach to AI Transparency
- Large Language Model Routing with Benchmark Datasets
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
- AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
Cited by
Related