A Survey on Large Language Models for Mathematical Reasoning
2025/06/10 by Peng Wang, Wang, Peng-Yuan, Liu, Tian-Shuo +18 · 19 citations
Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Text Readability and Simplification
paper · pdf · doi:10.48550/arxiv.2506.08446
Abstract
Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs) have achieved significant advances in this area. This survey examines the development of mathematical reasoning abilities in LLMs through two high-level cognitive phases: comprehension, where models gain mathematical understanding via diverse pretraining strategies, and answer generation, which has progressed from direct prediction to step-by-step Chain-of-Thought (CoT) reasoning. We review methods for enhancing mathematical reasoning, ranging from training-free prompting to fine-tuning approaches such as supervised fine-tuning and reinforcement learning, and discuss recent work on extended CoT and "test-time scaling". Despite notable progress, fundamental challenges remain in terms of capacity, efficiency, and generalization. To address these issues, we highlight promising research directions, including advanced pretraining and knowledge augmentation techniques, formal reasoning frameworks, and meta-generalization through principled learning paradigms. This survey tries to provide some insights for researchers interested in enhancing reasoning capabilities of LLMs and for those seeking to apply these techniques to other domains.
Citations
- Teaching Language Models to Reason with Tools
- TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- Reasoning Models Don't Always Say What They Think
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- xVerify: Efficient Answer Verifier for Reasoning Model Evaluations
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- ToRL: Scaling Tool-Integrated RL
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond
- Controlling Large Language Model with Latent Actions
- SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Self-rewarding correction for mathematical reasoning
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Reasoning with Reinforced Functional Token Tuning
- LIMO: Less is More for Reasoning
- Demystifying Long Chain-of-Thought Reasoning in LLMs
- Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
- s1: Simple test-time scaling
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- A Comprehensive Survey of the Lean 4 Theorem Prover: Architecture, Applications, and Advances
- LemmaHead: RAG Assisted Proof Generation Using Large Language Models
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
- A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
- Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
- Free Process Rewards without Process Labels
- SIKeD: Self-guided Iterative Knowledge Distillation for mathematical reasoning
- SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation
- Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
- Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
- ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
- LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
- LLM With Tools: A Survey
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- Planning In Natural Language Improves LLM Search For Code Generation
- On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
- Preserving Diversity in Supervised Fine-Tuning of Large Language Models
- Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- Critique-out-Loud Reward Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- The Llama 3 Herd of Models
- Recursive Introspection: Teaching Language Model Agents How to Self-Improve
- Qwen2 Technical Report
- Advancing Process Verification for Large Language Models via Tree-Based Preference Learning
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
- Lean Workbook: A large-scale Lean problem set formalized from natural language math problems
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
- BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
- Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
- Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
- SimPO: Simple Preference Optimization with a Reference-Free Reward
- JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models
- Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving
- MAmmoTH2: Scaling Instructions from the Web
- AlphaMath Almost Zero: Process Supervision without Process
- Iterative Reasoning Preference Optimization
- Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
- Token-level Direct Preference Optimization
- Rho-1: Not All Tokens Are What You Need
- Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
- Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
- Common 7B Language Models Already Possess Strong Math Capabilities
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning
- Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
- MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
- Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
- Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models
- Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering
- OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
- Chain-of-Thought Reasoning Without Prompting
- V-STaR: Training Verifiers for Self-Taught Reasoners
- InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Large Language Models for Mathematical Reasoning: Progresses and Challenges
- Augmenting Math Word Problems via Iterative Question Composing
- ReFT: Reasoning with Reinforced Fine-Tuning
- The Impact of Reasoning Step Length on Large Language Models
- MathPile: A Billion-Token-Scale Pretraining Corpus for Math
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- TinyGSM: achieving >80% on GSM8k with small language models
- Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
- Looped Transformers are Better at Learning Learning Algorithms
- Learning From Mistakes Makes LLM Better Reasoner
- TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models
- Why Can Large Language Models Generate Correct Chain-of-Thoughts?
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Llemma: An Open Language Model For Mathematics
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
- MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
- Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference
- Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Forward-Backward Reasoning in Large Language Models for Mathematical Verification
- Analyzing Chain-of-Thought Prompting in Large Language Models via Gradient-based Feature Attributions
- CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?
- LeanDojo: Theorem Proving with Retrieval-Augmented Language Models
- Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
- Boosting Language Models Reasoning with Chain-of-Knowledge Prompting
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Evaluating and Improving Tool-Augmented Computation-Intensive Math Reasoning
- Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
- Learning Multi-Step Reasoning by Solving Arithmetic Tasks
- Let's Verify Step by Step
- Leveraging Training Data in Few-Shot Prompting for Numerical Reasoning
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical Perspective
- Reasoning with Language Model is Planning with World Model
- Language Model Self-improvement by Reinforcement Learning Contemplation
- Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation
- TheoremQA: A Theorem-driven Question Answering dataset
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
- Why think step by step? Reasoning emerges from the locality of experience
- REFINER: Reasoning Feedback on Intermediate Representations
- A Survey of Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- MathPrompter: Mathematical Reasoning using Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters
- A Survey of Deep Learning for Mathematical Reasoning
- Large Language Models Are Reasoning Teachers
- Large Language Models are Better Reasoners with Self-Verification
- Teaching Small Language Models to Reason
- PAL: Program-aided Language Models
- Teaching Algorithmic Reasoning via In-context Learning
- Generating Sequences by Learning to Self-Correct
- Large Language Models Can Self-Improve
- Automatic Chain of Thought Prompting in Large Language Models
- Language Models are Multilingual Chain-of-Thought Reasoners
- Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions
- Solving Quantitative Reasoning Problems with Language Models
- Emergent Abilities of Large Language Models
- A Survey in Mathematical Language Processing
- Large Language Models are Zero-Shot Reasoners
- Chain of Thought Imitation with Procedure Cloning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Teaching language models to support answers with verified quotes
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- A General Language Assistant as a Laboratory for Alignment
- Training Verifiers to Solve Math Word Problems
- MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
- Program Synthesis with Large Language Models
- NaturalProofs: Mathematical Theorem Proving in Natural Language
- Are NLP Models really able to Solve Simple Math Word Problems?
- Measuring Mathematical Problem Solving With the MATH Dataset
- Measuring Massive Multitask Language Understanding
- Learning to summarize from human feedback
- A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers
- Fine-Tuning Language Models from Human Preferences
- MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
- The Gap of Semantic Parsing: A Survey on Automatic Math Word Problem Solvers
- Proximal Policy Optimization Algorithms
- Deep reinforcement learning from human preferences
- Thinking Fast and Slow with Deep Learning and Tree Search
- Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
- FastText.zip: Compressing text classification models
- Distilling the Knowledge in a Neural Network
- I.—COMPUTING MACHINERY AND INTELLIGENCE
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- The Expressive Power of Transformers with Chain of Thought
- Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
- Lila: A Unified Benchmark for Mathematical Reasoning
Cited by
Related