s1: Simple test-time scaling
2025/01/31 by Niklas Muennighoff, Zitong Yang, Muennighoff, Niklas +18 · 30 voices · 770 citations
Computer Science · Engineering · Mathematics · #Computer science #Fault Detection and Control Systems #Geology #Geometry #Mathematics #Neural Networks and Applications #Philosophy #Scaling #Simple (philosophy) #Test (biology)
paper · pdf · doi:10.48550/arxiv.2501.19393
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/01/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. We seek the simplest approach to achieve test-time scaling and strong reasoning performance. First, we curate a small dataset s1K of 1,000 questions paired with reasoning traces relying on three criteria we validate through ablations: difficulty, diversity, and quality. Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end. This can lead the model to double-check its answer, often fixing incorrect reasoning steps. After supervised finetuning the Qwen2.5-32B-Instruct language model on s1K and equipping it with budget forcing, our model s1-32B exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). Further, scaling s1-32B with budget forcing allows extrapolating beyond its performance without test-time intervention: from 50% to 57% on AIME24. Our model, data, and code are open-source at https://github.com/simplescaling/s1
Citations
Cited by
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
- SLPO: Scaling Latent Reasoning via a Surrogate Policy
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation
- Masked Visual Actions for Unified World Modeling
- Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
- LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- Length Penalties Make Chain-of-Thought Less Monitorable
- Compositional Diffusion with Guided Search for Long-Horizon Planning
- On-Policy Delta Distillation
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- How Inference Compute Shapes Frontier LLM Evaluation
- Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
- Base Models Know How to Reason, Thinking Models Learn When
- K2-Think: A Parameter-Efficient Reasoning System
- Humans Perceive Wrong Narratives from AI Reasoning Texts
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Hierarchical Reasoning Model
- Reinforcement Learning Teachers of Test Time Scaling
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- How malicious AI swarms can threaten democracy
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
- Reasoning Models Can Be Effective Without Thinking
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
- ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Reinforcement Learning via Self-Distillation
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
- Anthropocentric bias in language model evaluation
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Procedural Knowledge at Scale Improves Reasoning
- StAR: Segment Anything Reasoner
- Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems
- When Reasoning Meets Its Laws
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Meta-RL Induces Exploration in Language Agents
- Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Estimating problem difficulty without ground truth using Large Language Model comparisons
- AIR: Post-training Data Selection for Reasoning via Attention Head Influence
- State over Tokens: Characterizing the Role of Reasoning Tokens
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- Training Language Models to Use Prolog as a Tool
- Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- Verifier Threshold: An Efficient Test-Time Scaling Approach for Image Generation
- RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- CoRT: Code-integrated Reasoning within Thinking
- Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- ViT3: Unlocking Test-Time Training in Vision
- Knowledge Graph Augmented Large Language Models for Disease Prediction
- Mode-Conditioning Unlocks Superior Test-Time Scaling
- EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients
- G-KV: Decoding-Time KV Cache Eviction with Global Attention
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images
- Revisiting Generalization Across Difficulty Levels: It's Not So Easy
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
- Majority of the Bests: Improving Best-of-N via Bootstrapping
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Asking LLMs to Verify First is Almost Free Lunch
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- SafeRBench: Dissecting the Reasoning Safety of Large Language Models
- GPS: General Per-Sample Prompter
- Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
- BARD: budget-aware reasoning distillation
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Black-Box On-Policy Distillation of Large Language Models
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Rank-1 LoRAs Encode Interpretable Reasoning Signals
- Test-Time Iterative Error Correction for Efficient Diffusion Models
- Real-Time Reasoning Agents in Evolving Environments
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- Why Less is More (Sometimes): A Theory of Data Curation
- CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
ReMind: Understanding Deductive Code Reasoning in LLMs- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
- Reasoning Up the Instruction Ladder for Controllable Language Models
- AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
- Chain-of-Thought Hijacking
- e1: Learning Adaptive Control of Reasoning Effort
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- Multi-Agent Transactive Memory
- Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
- VISTA: A Test-Time Self-Improving Video Generation Agent
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Teaching Language Models to Reason with Tools
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
- Code-enabled language models can outperform reasoning models on diverse tasks
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
- Data-Centric Lessons To Improve Speech-Language Pretraining
- SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- Mapping Post-Training Forgetting in Language Models at Scale
- Online In-Context Distillation for Low-Resource Vision Language Models
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- LLM Agents Beyond Utility: An Open-Ended Perspective
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Budget-aware Test-time Scaling via Discriminative Verification
- Toward Reasoning-Centric Time-Series Analysis
- Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
- Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
- HoneyBee: Data Recipes for Vision-Language Reasoners
- Are Large Reasoning Models Interruptible?
- Demystifying Reinforcement Learning in Agentic Reasoning
- Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
- Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
- Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- Concise Reasoning in the Lens of Lagrangian Optimization
- Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
- Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
- Mitigating Overthinking through Reasoning Shaping
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
- DeepPrune: Parallel Scaling without Inter-trace Redundancy
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
- GCPO: When Contrast Fails, Go Gold
- Parallel Test-Time Scaling for Latent Reasoning Models
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
- CaRT: Teaching LLM Agents to Know When They Know Enough
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
- Test-Time Scaling of Reasoning Models for Machine Translation
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- Influence Functions for Efficient Data Selection in Reasoning
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
- GraphGhost: Tracing Structures Behind Large Language Models
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- Detecting Distillation Data from Reasoning Models
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- AlphaApollo: A System for Deep Agentic Reasoning
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- PatternKV: Flattening KV Representation Expands Quantization Headroom
- Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- Understanding the Role of Training Data in Test-Time Scaling
- Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- Generalized Parallel Scaling with Interdependent Generations
- Making, not Taking, the Best of N
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- Prompt Curriculum Learning for Efficient LLM Post-Training
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- On The Fragility of Benchmark Contamination Detection in Reasoning Models
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Entropy After ⟨
/Think ⟩ for reasoning model early exiting - Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
- OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
- RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
- SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- SpecExit: Accelerating Large Reasoning Model via Speculative Exit
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
- Poivre: Self-Refining Visual Pointing with Reinforcement Learning
- Evaluating Program Semantics Reasoning with Type Inference in System F
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Tracing Uncertainty in Language Model "Reasoning"
- HEART: Emotionally-driven test-time scaling of Language Models
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
- Variational Reasoning for Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- GRPO is Secretly a Process Reward Model
- Best-of-∞ -- Asymptotic Performance of Test-Time LLM Ensembling
- RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
- PolicyPad: Collaborative Prototyping of LLM Policies
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- SIM-CoT: Supervised Implicit Chain-of-Thought
- The Conductor and the Engine: A Path Towards Co-Designed Reasoning
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
- Proximal Supervised Fine-Tuning
- Reinforcement Learning on Pre-Training Data
- HyperAdapt: Simple High-Rank Adaptation
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Investigating Test-Time Scaling with Reranking for Machine Translation
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- LATTS: Locally Adaptive Test-Time Scaling
- When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition
- GrACE: A Generative Approach to Better Confidence Elicitation in Large Language Models
- Merge-of-Thought Distillation
- AdsQA: Towards Advertisement Video Understanding
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
- Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
- Reverse-Engineered Reasoning for Open-Ended Generation
- Chatbot To Help Patients Understand Their Health
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- Hunyuan-MT Technical Report
- Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
- Code Like Humans: A Multi-Agent Solution for Medical Coding
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- Throttling Web Agents Using Reasoning Gates
- Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Scaling Group Inference for Diverse and High-Quality Generation
- Dream 7B: Diffusion Large Language Models
- Deep Think with Confidence
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Lexical Hints of Accuracy in LLM Reasoning Chains
- Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- User-Assistant Bias in LLMs
- SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- Retrieval-augmented reasoning with lean language models
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
- Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
- Train Long, Think Short: Curriculum Learning for Efficient Reasoning
- Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- Sample-efficient LLM Optimization with Reset Replay
- LLM Unlearning Without an Expert Curated Dataset
- Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- LoRA is All You Need for Safety Alignment of Reasoning LLMs
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Enhancing Japanese Large Language Models with Reasoning Vectors
- The SMeL Test: A simple benchmark for media literacy in language models
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
- R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
- Hierarchical Budget Policy Optimization for Adaptive Reasoning
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
- LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
- ControlMed: Adding Reasoning Control to Medical Language Model
- Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
- Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language Models
- A2R2: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
- SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
- Diversity-Enhanced Reasoning for Subjective Questions
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
- Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
- A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
- A Neuroscience-Inspired Dual-Process Model of Compositional Generalization
- Does More Inference-Time Compute Really Help Robustness?
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Unlocking the Working Memory of Large Language Models for Latent Reasoning
- Natural-Language Agent Harnesses
- It's Not That Simple. An Analysis of Simple Test-Time Scaling
- Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
- CLARIFID: Improving Radiology Report Generation by Reinforcing Clinically Accurate Impressions and Enforcing Detailed Findings
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- Resa: Transparent Reasoning Models via SAEs
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- Vidar: Embodied Video Diffusion Model for Generalist Manipulation
- Agentar-DeepFinance-100K: A Large-Scale Financial Dataset via Systematic Chain-of-Thought Synthesis Optimization
- The Serial Scaling Hypothesis
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- ROC-n-reroll: How verifier imperfection affects test-time scaling
- ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
- SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning
- Learning to Reason Across Parallel Samples for LLM Reasoning
- A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- A Survey on Large Language Models for Mathematical Reasoning
- From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
- REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
- AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model
- Cycle Context Verification for In-Context Medical Image Segmentation
- BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Adaptive Termination for Multi-round Parallel Reasoning: An Universal Semantic Entropy-Guided Framework
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
- PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
- CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
- FEVO: Financial Knowledge Expansion and Reasoning Evolution for Large Language Models
- Steering Information Utility in Key-Value Memory for Language Model Post-Training
- Replacing thinking with tool usage enables reasoning in small language models
- Activation Steering for Chain-of-Thought Compression
- On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
- PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process
- Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
- APPO: Agentic Procedural Policy Optimization
- ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
- Controlling Thinking Speed in Reasoning Models
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
- Test-Time Scaling with Reflective Generative Model
- On Reasoning Strength Planning in Large Reasoning Models
- Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
- Reasoning as an Adaptive Defense for Safety
- ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy
- Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime
- Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
- Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models
- CyberV: Cybernetics for Test-time Scaling in Video Understanding
- ReasonBridge: Efficient Reasoning Transfer from Closed to Open-Source Language Models
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- Lost at the Beginning of Reasoning
- Through the Valley: Path to Effective Long CoT Training for Small Language Models
- APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
- Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning
- Ctrl-Z Sampling: Diffusion Sampling with Controlled Random Zigzag Explorations
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
- Scaling Speculative Decoding with Lookahead Reasoning
- KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
- From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
- ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- Less Data Less Tokens: Multilingual Unification Learning for Efficient Test-Time Reasoning in LLMs
- Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning
- AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- Bayesian Social Deduction with Graph-Informed Language Models
- AnyMAC: Cascading Flexible Multi-Agent Collaboration via Next-Agent Prediction
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
- DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
- Representation Consistency for Accurate and Coherent LLM Answer Aggregation
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
- Steering LLM Thinking with Budget Guidance
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
- QFFT, Question-Free Fine-Tuning for Adaptive Reasoning
- Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
- Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
- SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models
- SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows
- Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
- Simple Radiology VLLM Test-time Scaling with Thought Graph Traversal
- Eliciting Reasoning in Language Models with Cognitive Tools
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Efficient LLM Collaboration via Planning
- How Far Are We from Optimal Reasoning Efficiency?
- From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
- How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
- Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
- Code Execution as Grounded Supervision for LLM Reasoning
- ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization
- Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation
- Pareto Optimal Code Generation
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models
- CP-Bench: Evaluating Large Language Models for Constraint Modelling
- Scalable Chain of Thoughts via Elastic Reasoning
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- Unlocking Recursive Thinking of LLMs: Alignment via Refinement
- ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
- Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
- Accelerated Test-Time Scaling with Model-Free Speculative Sampling
- ProRefine: Inference-Time Prompt Refinement with Textual Feedback
- Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
- LLM-First Search: Self-Guided Exploration of the Solution Space
- Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps Translation
- Reason-to-Recommend: Using Interaction-of-Thought Reasoning to Enhance LLM Recommendation
- Kinetics: Rethinking Test-Time Scaling Laws
- ScaleRTL: Scaling LLMs with Reasoning Data and Test-Time Compute for Accurate RTL Code Generation
- Crosslingual Reasoning through Test-Time Scaling
- Reasoning or Overthinking: Evaluating Large Language Models on Financial Sentiment Analysis
- Training a Scientific Reasoning Model for Chemistry
- Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters
- Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
- CyclicReflex: Improving Reasoning Models via Cyclical Reflection Token Scheduling
- Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
- EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
- Long or short CoT? Investigating Instance-level Switch of Large Reasoning Models
- Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
- Natural, Artificial, and Human Intelligences
- AI Scientists Fail Without Strong Implementation Capability
- Incentivizing LLMs to Self-Verify Their Answers
- Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
- Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
- Scaling Textual Gradients via Sampling-Based Momentum
- Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
- Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
- A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
- TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence
- QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
- Grounded Reinforcement Learning for Visual Reasoning
- Table-R1: Inference-Time Scaling for Table Reasoning
- PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
- AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
- Learning Composable Chains-of-Thought
- Scaling Reasoning without Attention
- Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
- When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
- What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
- Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
- WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
- Scaling Offline RL via Efficient and Expressive Shortcut Models
- LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
- THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
- Who Reasons in the Large Language Models?
- Beyond Templates: Dynamic Adaptation of Reasoning Demonstrations via Feasibility-Aware Exploration
- Pretraining Language Models to Ponder in Continuous Space
- Self-Route: Automatic Mode Switching via Capability Estimation for Efficient Reasoning
- Thinker: Learning to Think Fast and Slow
- Can Past Experience Accelerate LLM Reasoning?
- FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
- Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law
- Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
- Faster and Better LLMs via Latency-Aware Test-Time Scaling
- Training-Free Multi-Step Audio Source Separation
- Incentivizing Strong Reasoning from Weak Supervision
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
- Concise Reasoning, Big Gains: Pruning Long Reasoning Trace with Difficulty-Aware Prompting
- Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
- Co-PatcheR: Collaborative Software Patching with Component(s)-specific Small Reasoning Models
- MMATH: A Multilingual Benchmark for Mathematical Reasoning
- Do Large Language Models (Really) Need Statistical Foundations?
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
- Test-Time Scaling of Diffusion Models via Noise Trajectory Search
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- Thought calibration: Efficient and confident test-time scaling
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- First Finish Search: Efficient Test-Time Scaling in Large Language Models
- Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
- The Real Barrier to LLM Agent Usability is Agentic ROI
- QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
- T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering
- Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
- Value-Guided Search for Efficient Chain-of-Thought Reasoning
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
- DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
- Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
- ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
- UFT: Unifying Supervised and Reinforcement Fine-Tuning
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language Models
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- MMaDA: Multimodal Large Diffusion Language Models
- How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
- When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
- When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- Large Language Models Implicitly Learn to See and Hear Just By Reading
- Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- Rank-K: Test-Time Reasoning for Listwise Reranking
- PRL: Prompts from Reinforcement Learning
- Activation-Guided Consensus Merging for Large Language Models
- DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
- Attend to Your Own Thoughts: Breaking the Barrier for Post-Training Quantization of Reasoning LLMs through the Lens of 1.58-Bit Quantization
- Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning
- General-Reasoner: Advancing LLM Reasoning Across All Domains
- AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound
- Think Only When You Need with Large Hybrid-Reasoning Models
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
- Efficient Agent Training for Computer Use
- Reasoning Models Better Express Their Confidence
- Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
- Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
- Are LLMs Better Formalizers than Solvers on Complex Problems?
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
- R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
- R3: Robust Rubric-Agnostic Reward Models
- Optimizing Anytime Reasoning via Budget Relative Policy Optimization
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- Thinking Short and Right Over Thinking Long: Serving LLM Reasoning Efficiently and Accurately
- DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
- Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
- LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
- HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
- VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
- Evaluating the Logical Reasoning Abilities of Large Reasoning Models
- J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
- SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
- Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
- Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
- Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
- Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- SOD: Step-wise On-policy Distillation for Small Language Model Agents
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- Adaptive Social Learning via Mode Policy Optimization for Language Agents
- Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
- Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
- Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers
- Llama-Nemotron: Efficient Reasoning Models
- Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
- Recursive Multi-Agent Systems
- Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
- Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
- Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
- Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
- How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
- TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
- Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
- Architecting Trust in Artificial Epistemic Agents
- ParamMem: Augmenting Language Agents with Parametric Reflective Memory
- When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
- Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
- Benchmark Test-Time Scaling of General LLM Agents
- Training Large Reasoning Models Efficiently via Progressive Thought Encoding
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
- Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
- Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key
- Gradient Extrapolation-Based Policy Optimization
- When Less is Enough: Efficient Inference via Collaborative Reasoning
- Computational Reasoning of Large Language Models
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
- AutoJudge: Judge Decoding Without Manual Annotation
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
- Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
- AI Achieves a Perfect LSAT Score
- An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
- DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
- Concise Reasoning via Reinforcement Learning
- Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
- Rethinking Reflection in Pre-Training
- Frontier AI's Impact on the Cybersecurity Landscape
- SplitReason: Learning To Offload Reasoning
- AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
- Process Reward Models That Think
- Synergizing RAG and Reasoning: A Systematic Review
- Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
- Dynamic Early Exit in Reasoning Models
- DianJin-R1: Evaluating and Enhancing Financial Reasoning in Large Language Models
- Tina: Tiny Reasoning Models via LoRA
- Compass-V2 Technical Report
- Learning to Reason under Off-Policy Guidance
- Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
- LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
- Contemplative Agent
- a1: Steep Test-time Scaling Law via Environment Augmented Generation
- ReasoningV: Efficient Verilog Code Generation with Adaptive Hybrid Reasoning Model
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
- Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading
- Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
- Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
- Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
- Efficient Reasoning Models: A Survey
- Agent-Q: Fine-Tuning Large Language Models for Quantum Circuit Generation and Optimization
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
- Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
- MIEB: Massive Image Embedding Benchmark
- Weight Ensembling Improves Reasoning in Language Models
- Guiding Reasoning in Small Language Models with LLM Assistance
- Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning
- Exploring the System 1 Thinking Capability of Large Reasoning Models
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
- Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
- Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models
- SEAL: Steerable Reasoning Calibration of Large Language Models for Free
- Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
- Feedback neural network [wikipedia]
- Reasoning model [wikipedia]
Discussions
- This paper is wild - a Stanford team shows the simplest way to make an open LLM into a reasoning model They used just 1,000 carefully curated reasoning examples & a trick where if the model tries to [bsky, 214 points, 7 comments]
- « appending "Wait" multiple times to the model's generation » is our current most likely path to AGI :) See the fresh arxiv.org/abs/2501.19393 by Niklas Muennighoff et al. [bsky, 62 points, 0 comments]
- s1: Simple inference-time scaling This is a simple small-scale replication of inference-time scaling It was cheap: 16xH100 for 26 minutes (so what, ~$6?) It replicates inference-time scaling using [bsky, 29 points, 1 comments]
- - Link to the paper: arxiv.org/abs/2501.19393 - My reasoning LLM article (they use methods 1 and 4): magazine.sebastianraschka.com/p/understand... - The s1 GitHub repo: github.com/simplescalin... [bsky, 8 points, 0 comments]
- S1: Simple Test-Time Scaling [hn, 3 points, 0 comments]
- stumbled upon this amazing paper: arxiv.org/abs/2501.19393 Test-Time Scaling Simplified: authors used budget forcing and careful data selection to #SFT s1-32B model to enhance #reasoning & math perfor [bsky, 3 points, 0 comments]
- Test-time scaling new approach: extra test-time compute improves LLM reasoning [hn, 2 points, 0 comments]
- Hilarious: Just injecting a little self doubt by appending a "Wait" impressively improves model performance. Next step: Metacognitive processes baked into the network architecture? [bsky, 2 points, 0 comments]
- I worry about such click baiting titles: AI researchers at Stanford and the University of Washington were able to train an AI “reasoning” model for under $50. This was prompted by this Arxiv submiss [bsky, 2 points, 0 comments]
- s1: Simple Test-Time Scaling [hn, 2 points, 0 comments]
- In case you’re wondering, they also tested the performance of “Wait” vs. “Alternatively” and “Hmm” The future of computer science, right here! arxiv.org/pdf/2501.19393 [bsky, 2 points, 0 comments]
- This podcast is what made me realize how much behavior can change post-training twimlai.com/podcast/twim... [bsky, 1 points, 0 comments]
- This is pretty amazing stuff. If you get a chance, check out the original paper, which includes this very fun table testing the efficacy of extending "thinking" time. arxiv.org/pdf/2501.19393 [bsky, 1 points, 0 comments]
- This one simple trick applies more compute at inference to improve model reasoning by adjusting output duration. The work details a question set with reasoning traces and finetunes a model to alter it [bsky, 1 points, 0 comments]
- arxiv.org/abs/2501.19393 [bsky, 1 points, 0 comments]
- what if all that's needed a sort of SFT as done in the s1: Simple test-time scaling paper arxiv.org/abs/2501.193... but for interdisciplinary thinking? why expect LLMs to give answers that are exceedi [bsky, 1 points, 0 comments]
- A hilariously simple repro of OpenAI's test-time scaling paradigm called "Budget Scaling": end the thinking when your token budget is met, or append "Wait" to the model's generation to keep thinking, [bsky, 1 points, 0 comments]
- Cue the next DeepSeek headline 😂 Impressive work from Stanford team! arxiv.org/pdf/2501.19393 [bsky, 1 points, 0 comments]
- Fei-Fei Li's lab released a new #AIreasoning model (called s1) that performs on par with the big ones, but cost only $50 to train and uses 1,000 samples. @techcrunch.com: techcrunch.com/2025/02/05/r. [bsky, 1 points, 0 comments]
- Many test-time-compute papers show why this is hard. In an extreme version, just appending "Wait" whenever the model tried to stop raised the benchmark. Often, just forcing more reasoning does work: a [bsky, 1 points, 1 comments]
- Hey, if you get a minute, this paper on AI pre-deployment-tuning is fucking fire. arxiv.org/pdf/2501.19393 [bsky, 0 points, 0 comments]
- Interesting paper from Niklas Muennighoff, Zitong Yang et al. from Stanford arxiv.org/abs/2501.19393 "Thus, we ask: what is the simplest approach to achieve both test-time scaling and strong reasoni [bsky, 0 points, 0 comments]
- s1: Simple test-time scaling #llms #reasoning #wait #s1 #thinking [bsky, 0 points, 0 comments]
- S1, a 50 dollar o1 competitor arxiv.org/pdf/2501.19393 Looks like @profgalloway.com is right about the OpenAI bubble. [bsky, 0 points, 0 comments]
- Yet another LLM. s1 is developed by Washington and Stanford universities and it's free. arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- Adding "Wait" multiple times increased accuracy even further (see image). If you've had any experience using DeepSeek, then you've probably already seen this behaviour. This may explain why. Paper [bsky, 0 points, 0 comments]
- Stanford / UW Team showcases accuracy improvements in reasoning through test-time scaling, high quality data, and brute forcing reasoning backtracking on Qwen 32B fine tune arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- Simple test-time scaling [Muennighoff+, 2025] The authors reproduced the test-time scaling curve by fine-tuning Qwen2.5-32B with s1K, a set of reasoning traces generated by Gemini. They controlled the [bsky, 0 points, 0 comments]
- [2501.19393] s1: Simple test-time scaling arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
Related