rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
2025/01/08 by Xinyu Guan, Li Lyna Zhang, Guan, Xinyu +15 · 11 voices · 156 citations
Computer Science · Psychology · #Intelligent Tutoring Systems and Adaptive Learning #Machine Learning and Data Classification #Mathematics education #Mathematics, Computing, and Information Processing #Psychology
paper · pdf · doi:10.48550/arxiv.2501.04519
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/01/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models. rStar-Math achieves this by exercising "deep thinking" through Monte Carlo Tree Search (MCTS), where a math policy SLM performs test-time search guided by an SLM-based process reward model. rStar-Math introduces three innovations to tackle the challenges in training the two SLMs: (1) a novel code-augmented CoT data sythesis method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories used to train the policy SLM; (2) a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective process preference model (PPM); (3) a self-evolution recipe in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities. Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems, rStar-Math boosts SLMs' math reasoning to state-of-the-art levels. On the MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0% and Phi3-mini-3.8B from 41.4% to 86.4%, surpassing o1-preview by +4.5% and +0.9%. On the USA Math Olympiad (AIME), rStar-Math solves an average of 53.3% (8/15) of problems, ranking among the top 20% the brightest high school math students. Code and data will be available at https://github.com/microsoft/rStar.
Cited by
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
- M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- PyFi: Toward Pyramid-like Financial Image Understanding for VLMs via Adversarial Agents
- Evolutionary System 2 Reasoning: An Empirical Proof
- CoRT: Code-integrated Reasoning within Thinking
- CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code Generation
- Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
- Boosting In-Silicon Directed Evolution with Fine-Tuned Protein Language Model and Tree Search
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- Reasoning Planning for Language Models
- SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- MedRule-KG: A Knowledge-Graph--Steered Scaffold for Mathematical Reasoning with a Lightweight Verifier
- MASPRM: Multi-Agent System Process Reward Model
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Teaching Language Models to Reason with Tools
- Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
- LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning
- Budget-aware Test-time Scaling via Discriminative Verification
- Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
- A Survey on Agentic Multimodal Large Language Models
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- Pinpointing crucial steps: Attribution-based Credit Assignment for Verifiable Reinforcement Learning
- A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
- Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
- Meta-R1: Empowering Large Reasoning Models with Metacognition
- MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- rStar2-Agent: Agentic Reasoning Technical Report
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
- Sample-efficient LLM Optimization with Reset Replay
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
- Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis
- Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
- Decoupling Knowledge and Reasoning in LLMs: An Exploration Using Cognitive Dual-System Theory
- AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
- Know What You Don't Know: Uncertainty Calibration of Process Reward Models
- Agentar-DeepFinance-100K: A Large-Scale Financial Dataset via Systematic Chain-of-Thought Synthesis Optimization
- Med-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- A Survey on Large Language Models for Mathematical Reasoning
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
- AI-Powered Math Tutoring: Platform for Personalized and Adaptive Education
- EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
- Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
- From Language to Logic: A Bi-Level Framework for Structured Reasoning
- Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS
- Replacing thinking with tool usage enables reasoning in small language models
- ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
- Test-Time Scaling with Reflective Generative Model
- Frontiers of Generative AI for Network Optimization: Theories, Limits, and Visions
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- Distilling Tool Knowledge into Language Models via Back-Translated Traces
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
- The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
- RL for Reasoning by Adaptively Revealing Rationales
- DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
- KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
- ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
- MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Spurious Rewards: Rethinking Training Signals in RLVR
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- Unlocking Recursive Thinking of LLMs: Alignment via Refinement
- DynamicMind: A Tri-Mode Thinking System for Large Language Models
- SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
- Structured Pruning for Diverse Best-of-N Reasoning Optimization
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
- One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
- Table-R1: Inference-Time Scaling for Table Reasoning
- MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration
- RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
- R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
- Can Past Experience Accelerate LLM Reasoning?
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- Faster and Better LLMs via Latency-Aware Test-Time Scaling
- REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
- Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
- MMATH: A Multilingual Benchmark for Mathematical Reasoning
- VeriThinker: Learning to Verify Makes Reasoning Model Efficient
- Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- RaDeR: Reasoning-aware Dense Retrieval Models
- Learning to Reason via Mixture-of-Thought for Logical Reasoning
- RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
- TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
- From Reasoning to Code: GRPO Optimization for Underrepresented Languages
- DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
- Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
- TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models
- Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
- Chain-of-Thought Tokens are Computer Program Variables
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- Self-Improving Language Models with Bidirectional Evolutionary Search
- Phi-4-reasoning Technical Report
- Self-Verified Distillation: Your Language Model Is Secretly Its Own Synthetic Data Pipeline
- Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
- R3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
- Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
- Can Post-Training Transform LLMs into Causal Reasoners?
- IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
- OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents
- Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
- MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
Discussions
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking [hn, 39 points, 7 comments]
- Lots of other very clever stuff in the paper: arxiv.org/pdf/2501.045... [bsky, 17 points, 1 comments]
- rStar-Math takes qwen2.5 7b & 1.5b as well as qhi3 3.8b and fine tunes them for math they’re able to exceed o1-preview on math benchmarks (with the 7B) the magic sauce seems to be in co-evolving the [bsky, 9 points, 1 comments]
- Paper: arxiv.org/abs/2501.04519 @microsoft.com Research [bsky, 2 points, 0 comments]
- rStar-Math is the most elegant application of MCTS in RL I've seen arxiv.org/abs/2501.04519 [bsky, 2 points, 0 comments]
- Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking [hn, 1 points, 0 comments]
- 🧪🧪 and for our #AI research of the week... “rStar-Math” introduces a framework showcasing that smaller language models can achieve state-of-the-art math reasoning capabilities through iterative sel [bsky, 0 points, 1 comments]
- Paper on arXiv: arxiv.org/abs/2501.04519 [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2501.04519 [bsky, 0 points, 0 comments]
- Well, by doing the same thing with 'small' LLMs, they're at least being marginally less wasteful arxiv.org/pdf/2501.04519 [bsky, 0 points, 1 comments]
- arxiv.org/abs/2501.04519 Das bedeutet, dass man für PhD Level reasoning keine unvorstellbaren compute-cluster benötigt. Auch weiterhin wird Intelligenz günstiger. [bsky, 0 points, 0 comments]
Related