MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
2023/09/21 by Longhui Yu, Weisen Jiang, Yu, Longhui +17 · 1 voice · 197 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2309.12284
arxiv published 2023/09/21 · arxiv updated 2024/05/03
Abstract
Large language models (LLMs) have pushed the limits of natural language understanding and exhibited excellent problem-solving ability. Despite the great success, most existing open-source LLMs (e.g., LLaMA-2) are still far away from satisfactory for solving mathematical problem due to the complex reasoning procedures. To bridge this gap, we propose MetaMath, a fine-tuned language model that specializes in mathematical reasoning. Specifically, we start by bootstrapping mathematical questions by rewriting the question from multiple perspectives without extra knowledge, which results in a new dataset called MetaMathQA. Then we fine-tune the LLaMA-2 models on MetaMathQA. Experimental results on two popular benchmarks (i.e., GSM8K and MATH) for mathematical reasoning demonstrate that MetaMath outperforms a suite of open-source LLMs by a significant margin. Our MetaMath-7B model achieves 66.4% on GSM8K and 19.4% on MATH, exceeding the state-of-the-art models of the same size by 11.5% and 8.7%. Particularly, MetaMath-70B achieves an accuracy of 82.3% on GSM8K, slightly better than GPT-3.5-Turbo. We release all the MetaMathQA dataset, the MetaMath models with different model sizes and the training code for public use.
Cited by
- The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- Greedy Alignment Principle for Optimizer Selection
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
- Parameter-Efficient Subspace Optimization for LLM Fine-Tuning
- Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
- NVIDIA Nemotron Nano V2 VL
- Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- Autodata: An agentic data scientist to create high quality synthetic data
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- DictPFL: Efficient and Private Federated Learning on Encrypted Gradients
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning
- FineVision: Open Data Is All You Need
- StreamingThinker: Large Language Models Can Think While Reading
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Interpretability Framework for LLMs in Undergraduate Calculus
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
- CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- A Survey on Evaluation of Large Language Models
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- AMiD: Knowledge Distillation for LLMs with α-mixture Assistant Distribution
- Skill-Targeted Adaptive Training
- StelLA: Subspace Learning in Low-rank Adaptation using Stiefel Manifold
- Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Mid-Training of Large Language Models: A Survey
- KaVa: Latent Reasoning via Compressed KV-Cache Distillation
- POME: Post Optimization Model Edit via Muon-style Projection
- Making Mathematical Reasoning Adaptive
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- TaskCraft: Automated Generation of Agentic Tasks
- EvolProver: Advancing Automated Theorem Proving by Evolving Formalized Problems via Symmetry and Difficulty
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
- Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
- Communication-Efficient and Accurate Approach for Aggregation in Federated Low-Rank Adaptation
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Plan before Solving: Problem-Aware Strategy Routing for Mathematical Reasoning with LLMs
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
- Delta Activations: A Representation for Finetuned Large Language Models
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
- Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
- Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
- Sample-efficient LLM Optimization with Reset Replay
- Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
- CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
- MathDuels: Evaluating LLMs as Problem Posers and Solvers
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- The Primacy of Magnitude in Low-Rank Adaptation
- BlueLM-2.5-3B Technical Report
- Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
- ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
- Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- Improving Large Language Models with Concept-Aware Fine-Tuning
- EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
- Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
- PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
- Distilling Tool Knowledge into Language Models via Back-Translated Traces
- Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
- SPARQ: Synthetic Problem Generation for Reasoning via Quality-Diversity Algorithms
- SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- TreeRPO: Tree Relative Policy Optimization
- Crosslingual Reasoning through Test-Time Scaling
- CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
- PoLAR: Polar-Decomposed Low-Rank Adapter Representation
- DiaBlo: Diagonal Blocks Are Sufficient For Finetuning
- Is Extending Modality The Right Path Towards Omni-Modality?
- MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- Uni-LoRA: One Vector is All You Need
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought
- AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning
- Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
- How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
- Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
- Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
- FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
- Towards Better Instruction Following Retrieval Models
- Can Large Reasoning Models Self-Train?
- Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
- MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
- GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
- Behavior Injection: Preparing Language Models for Reinforcement Learning
- Knowledge Grafting of Large Language Models
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
- Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
- Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
- Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
- Optimal Policy Minimum Bayesian Risk
- HOFT: Householder Orthogonal Fine-tuning
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
- DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
- Learning to Rank Chain-of-Thought: Using a Small Model
- Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
- Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
- Context-Free Synthetic Data Mitigates Forgetting
- ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
- Learnware of Language Models: Specialized Small Language Models Can Do Big
- MARGE: Improving Math Reasoning for LLMs with Guided Exploration
- AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating Projections
- HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation
- Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning
- Critique-Guided Distillation for Robust Reasoning via Refinement
- CLT and Edgeworth Expansion for m-out-of-n Bootstrap Estimators of The Studentized Median
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models
- PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
- AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
- SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
- On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation
- Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
- DeepCritic: Deliberate Critique with Large Language Models
- FineScope : Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation
- Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
- Predictable GRPO: A Closed-Form Model of Training Dynamics
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- DEL: Digit Entropy Loss for Numerical Learning of Large Language Models
- Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
- Sensitivity-Positional Co-Localization in GQA Transformers
- Multi-Token Prediction via Self-Distillation
- DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
- Frontier AI's Impact on the Cybersecurity Landscape
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Efficient Reasoning Models: A Survey
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
- Weight Ensembling Improves Reasoning in Language Models
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- Fine-tuning a Large Language Model for Automating Computational Fluid Dynamics Simulations
Discussions
Related