A Primer in Post-Training Reasoning Data: What We Know About How It Works
2026/06/01 by Yaoming Li, Guangxiang Zhao, Qilong Shi +3 · 3 voices
Computer Science · #cs.CL #cs.AI
paper · pdf
Abstract
Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.
Citations
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- A Survey on LLM Mid-Training
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Online Rubrics Elicitation from Pairwise Comparisons
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
- GRPO is Secretly a Process Reward Model
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- One Token to Fool LLM-as-a-Judge
- OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Spurious Rewards: Rethinking Training Signals in RLVR
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
- Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
- TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
- R3: Robust Rubric-Agnostic Reward Models
- AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
- Distillation Scaling Laws
- LIMO: Less is More for Reasoning
- s1: Simple test-time scaling
- Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
- The Lessons of Developing Process Reward Models in Mathematical Reasoning
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
- Training Software Engineering Agents and Verifiers with SWE-Gym
- ProcessBench: Identifying Process Errors in Mathematical Reasoning
- The BrowserGym Ecosystem for Web Agent Research
- A Note on Shumailov et al. (2024): `AI Models Collapse When Trained on Recursively Generated Data'
- Language agents achieve superhuman synthesis of scientific knowledge
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- Autonomous Tester Agent Benchmark
- Mind2Web: Towards a Generalist Agent for the Web
- Let's Verify Step by Step
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Gorilla: Large Language Model Connected with Massive APIs
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Scaling Laws for Reward Model Overoptimization
- Training language models to follow instructions with human feedback
- Training Verifiers to Solve Math Word Problems
- ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Finetuned Language Models Are Zero-Shot Learners
- FinQA: A Dataset of Numerical Reasoning over Financial Data
- MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
- Evaluating Large Language Models Trained on Code
- Measuring Coding Challenge Competence With APPS
- TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
- A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
- CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
- Measuring Mathematical Problem Solving With the MATH Dataset
- Fact or Fiction: Verifying Scientific Claims
- PubMedQA: A Dataset for Biomedical Research Question Answering
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Discussions
Related