s1: Simple test-time scaling
2025/01/31 by Niklas Muennighoff, Zitong Yang, Muennighoff, Niklas +18 · 30 voices · 378 citations
Engineering · Computer Science · #Fault Detection and Control Systems #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.2501.19393
Abstract
Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. We seek the simplest approach to achieve test-time scaling and strong reasoning performance. First, we curate a small dataset s1K of 1,000 questions paired with reasoning traces relying on three criteria we validate through ablations: difficulty, diversity, and quality. Second, we develop budget forcing to control test-time compute by forcefully terminating the model's thinking process or lengthening it by appending "Wait" multiple times to the model's generation when it tries to end. This can lead the model to double-check its answer, often fixing incorrect reasoning steps. After supervised finetuning the Qwen2.5-32B-Instruct language model on s1K and equipping it with budget forcing, our model s1-32B exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). Further, scaling s1-32B with budget forcing allows extrapolating beyond its performance without test-time intervention: from 50% to 57% on AIME24. Our model, data, and code are open-source at https://github.com/simplescaling/s1
Citations
Cited by
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
- SLPO: Scaling Latent Reasoning via a Surrogate Policy
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation
- Masked Visual Actions for Unified World Modeling
- Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
- LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- Length Penalties Make Chain-of-Thought Less Monitorable
- Compositional Diffusion with Guided Search for Long-Horizon Planning
- On-Policy Delta Distillation
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- How Inference Compute Shapes Frontier LLM Evaluation
- Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
- Base Models Know How to Reason, Thinking Models Learn When
- K2-Think: A Parameter-Efficient Reasoning System
- Humans Perceive Wrong Narratives from AI Reasoning Texts
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Hierarchical Reasoning Model
- Reinforcement Learning Teachers of Test Time Scaling
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- How malicious AI swarms can threaten democracy
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
- Reasoning Models Can Be Effective Without Thinking
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
- ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Reinforcement Learning via Self-Distillation
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
- Anthropocentric bias in language model evaluation
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Procedural Knowledge at Scale Improves Reasoning
- StAR: Segment Anything Reasoner
- Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems
- When Reasoning Meets Its Laws
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Meta-RL Induces Exploration in Language Agents
- Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- Estimating problem difficulty without ground truth using Large Language Model comparisons
- AIR: Post-training Data Selection for Reasoning via Attention Head Influence
- State over Tokens: Characterizing the Role of Reasoning Tokens
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
- Training Language Models to Use Prolog as a Tool
- Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- An Efficient Test-Time Scaling Approach for Image Generation
- RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists
- Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- ViT3: Unlocking Test-Time Training in Vision
- Knowledge Graph Augmented Large Language Models for Disease Prediction
- Mode-Conditioning Unlocks Superior Test-Time Scaling
- EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients
- G-KV: Decoding-Time KV Cache Eviction with Global Attention
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- PathReasoning: A multimodal reasoning agent for query-based ROI navigation on whole-slide images
- Revisiting Generalization Across Difficulty Levels: It's Not So Easy
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
- Majority of the Bests: Improving Best-of-N via Bootstrapping
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- Asking LLMs to Verify First is Almost Free Lunch
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- SafeRBench: A Comprehensive Benchmark for Safety Assessment in Large Reasoning Models
- GPS: General Per-Sample Prompter
- Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
- BARD: budget-aware reasoning distillation
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Black-Box On-Policy Distillation of Large Language Models
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Rank-1 LoRAs Encode Interpretable Reasoning Signals
- Test-Time Iterative Error Correction for Efficient Diffusion Models
- Real-Time Reasoning Agents in Evolving Environments
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- Why Less is More (Sometimes): A Theory of Data Curation
- CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
ReMind: Understanding Deductive Code Reasoning in LLMs- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
- Reasoning Up the Instruction Ladder for Controllable Language Models
- AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
- Chain-of-Thought Hijacking
- e1: Learning Adaptive Control of Reasoning Effort
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- SymCode: A Neurosymbolic Approach to Mathematical Reasoning via Verifiable Code Generation
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- Multi-Agent Transactive Memory
- Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
- RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
- VISTA: A Test-Time Self-Improving Video Generation Agent
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
- Teaching Language Models to Reason with Tools
- RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
- Code-enabled language models can outperform reasoning models on diverse tasks
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
- Data-Centric Lessons To Improve Speech-Language Pretraining
- SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- Test-time Verification via Optimal Transport: Coverage, ROC, & Sub-optimality
- Mapping Post-Training Forgetting in Language Models at Scale
- Online In-Context Distillation for Low-Resource Vision Language Models
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Visual Autoregressive Models Beat Diffusion Models on Inference Time Scaling
- LLM Agents Beyond Utility: An Open-Ended Perspective
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Budget-aware Test-time Scaling via Discriminative Verification
- Toward Reasoning-Centric Time-Series Analysis
- Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
- Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
- HoneyBee: Data Recipes for Vision-Language Reasoners
- Are Large Reasoning Models Interruptible?
- Demystifying Reinforcement Learning in Agentic Reasoning
- Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
- EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
- Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
- Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
- Trace Length is a Simple Uncertainty Signal in Reasoning Models
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- Concise Reasoning in the Lens of Lagrangian Optimization
- Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning
- Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
- Mitigating Overthinking through Reasoning Shaping
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
- DeepPrune: Parallel Scaling without Inter-trace Redundancy
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
- GCPO: When Contrast Fails, Go Gold
- Parallel Test-Time Scaling for Latent Reasoning Models
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
- CaRT: Teaching LLM Agents to Know When They Know Enough
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
- Test-Time Scaling of Reasoning Models for Machine Translation
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectory?
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- Influence Functions for Efficient Data Selection in Reasoning
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
- GraphGhost: Tracing Structures Behind Large Language Models
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- Detecting Distillation Data from Reasoning Models
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- AlphaApollo: Orchestrating Foundation Models and Professional Tools into a Self-Evolving System for Deep Agentic Reasoning
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- PatternKV: Flattening KV Representation Expands Quantization Headroom
- Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- Understanding the Role of Training Data in Test-Time Scaling
- Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
- Generalized Parallel Scaling with Interdependent Generations
- Making, not Taking, the Best of N
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- Prompt Curriculum Learning for Efficient LLM Post-Training
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- On The Fragility of Benchmark Contamination Detection in Reasoning Models
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Entropy After ⟨
/Think ⟩ for reasoning model early exiting - Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
- OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
- RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
- Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
- SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
- ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- SpecExit: Accelerating Large Reasoning Model via Speculative Exit
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step
- Poivre: Self-Refining Visual Pointing with Reinforcement Learning
- Evaluating Program Semantics Reasoning with Type Inference in System F
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Tracing Uncertainty in Language Model "Reasoning"
- HEART: Emotionally-driven test-time scaling of Language Models
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
- Variational Reasoning for Language Models
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- GRPO is Secretly a Process Reward Model
- Best-of-∞ -- Asymptotic Performance of Test-Time LLM Ensembling
- RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
- PolicyPad: Collaborative Prototyping of LLM Policies
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- SIM-CoT: Supervised Implicit Chain-of-Thought
- The Conductor and the Engine: A Path Towards Co-Designed Reasoning
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
- Proximal Supervised Fine-Tuning
- Reinforcement Learning on Pre-Training Data
- HyperAdapt: Simple High-Rank Adaptation
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Investigating Test-Time Scaling with Reranking for Machine Translation
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- ATTS: Asynchronous Test-Time Scaling via Conformal Prediction
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- LATTS: Locally Adaptive Test-Time Scaling
- When Inverse Data Outperforms: Exploring the Pitfalls of Mixed Data in Multi-Stage Fine-Tuning
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- Using LLMs for Late Multimodal Sensor Fusion for Activity Recognition
- GrACE: A Generative Approach to Better Confidence Elicitation in Large Language Models
- Merge-of-Thought Distillation
- AdsQA: Towards Advertisement Video Understanding
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
- Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
- Reverse-Engineered Reasoning for Open-Ended Generation
- Chatbot To Help Patients Understand Their Health
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- Hunyuan-MT Technical Report
- Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
- Code Like Humans: A Multi-Agent Solution for Medical Coding
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
- Throttling Web Agents Using Reasoning Gates
- Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Scaling Group Inference for Diverse and High-Quality Generation
- Dream 7B: Diffusion Large Language Models
- Deep Think with Confidence
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Lexical Hints of Accuracy in LLM Reasoning Chains
- Input-Time Scaling
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- User-Assistant Bias in LLMs
- SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- Retrieval-augmented reasoning with lean language models
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- mSCoRe: a Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
- PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
- Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
- Train Long, Think Short: Curriculum Learning for Efficient Reasoning
- Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- Sample-efficient LLM Optimization with Reset Replay
- LLM Unlearning Without an Expert Curated Dataset
- Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Enhancing Japanese Large Language Models with Reasoning Vectors
- The SMeL Test: A simple benchmark for media literacy in language models
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
- R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
- CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
- Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
- ControlMed: Adding Reasoning Control to Medical Language Model
- Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
- A2R2: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
- SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
- Diversity-Enhanced Reasoning for Subjective Questions
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
- Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
- A Toolbox, Not a Hammer -- Multi-TAG: Scaling Math Reasoning with Multi-Tool Aggregation
- A Neuroscience-Inspired Dual-Process Model of Compositional Generalization
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Unlocking the Working Memory of Large Language Models for Latent Reasoning
- Natural-Language Agent Harnesses
- Feedback neural network [wikipedia]
- Reasoning model [wikipedia]
Discussions
- This paper is wild - a Stanford team shows the simplest way to make an open LLM into a reasoning model They used just 1,000 carefully curated reasoning examples & a trick where if the model tries to [bsky, 214 points, 7 comments]
- « appending "Wait" multiple times to the model's generation » is our current most likely path to AGI :) See the fresh arxiv.org/abs/2501.19393 by Niklas Muennighoff et al. [bsky, 62 points, 0 comments]
- s1: Simple inference-time scaling This is a simple small-scale replication of inference-time scaling It was cheap: 16xH100 for 26 minutes (so what, ~$6?) It replicates inference-time scaling using [bsky, 29 points, 1 comments]
- - Link to the paper: arxiv.org/abs/2501.19393 - My reasoning LLM article (they use methods 1 and 4): magazine.sebastianraschka.com/p/understand... - The s1 GitHub repo: github.com/simplescalin... [bsky, 8 points, 0 comments]
- S1: Simple Test-Time Scaling [hn, 3 points, 0 comments]
- stumbled upon this amazing paper: arxiv.org/abs/2501.19393 Test-Time Scaling Simplified: authors used budget forcing and careful data selection to #SFT s1-32B model to enhance #reasoning & math perfor [bsky, 3 points, 0 comments]
- Test-time scaling new approach: extra test-time compute improves LLM reasoning [hn, 2 points, 0 comments]
- Hilarious: Just injecting a little self doubt by appending a "Wait" impressively improves model performance. Next step: Metacognitive processes baked into the network architecture? [bsky, 2 points, 0 comments]
- I worry about such click baiting titles: AI researchers at Stanford and the University of Washington were able to train an AI “reasoning” model for under $50. This was prompted by this Arxiv submiss [bsky, 2 points, 0 comments]
- s1: Simple Test-Time Scaling [hn, 2 points, 0 comments]
- In case you’re wondering, they also tested the performance of “Wait” vs. “Alternatively” and “Hmm” The future of computer science, right here! arxiv.org/pdf/2501.19393 [bsky, 2 points, 0 comments]
- This podcast is what made me realize how much behavior can change post-training twimlai.com/podcast/twim... [bsky, 1 points, 0 comments]
- This is pretty amazing stuff. If you get a chance, check out the original paper, which includes this very fun table testing the efficacy of extending "thinking" time. arxiv.org/pdf/2501.19393 [bsky, 1 points, 0 comments]
- This one simple trick applies more compute at inference to improve model reasoning by adjusting output duration. The work details a question set with reasoning traces and finetunes a model to alter it [bsky, 1 points, 0 comments]
- arxiv.org/abs/2501.19393 [bsky, 1 points, 0 comments]
- what if all that's needed a sort of SFT as done in the s1: Simple test-time scaling paper arxiv.org/abs/2501.193... but for interdisciplinary thinking? why expect LLMs to give answers that are exceedi [bsky, 1 points, 0 comments]
- A hilariously simple repro of OpenAI's test-time scaling paradigm called "Budget Scaling": end the thinking when your token budget is met, or append "Wait" to the model's generation to keep thinking, [bsky, 1 points, 0 comments]
- Cue the next DeepSeek headline 😂 Impressive work from Stanford team! arxiv.org/pdf/2501.19393 [bsky, 1 points, 0 comments]
- Fei-Fei Li's lab released a new #AIreasoning model (called s1) that performs on par with the big ones, but cost only $50 to train and uses 1,000 samples. @techcrunch.com: techcrunch.com/2025/02/05/r. [bsky, 1 points, 0 comments]
- Many test-time-compute papers show why this is hard. In an extreme version, just appending "Wait" whenever the model tried to stop raised the benchmark. Often, just forcing more reasoning does work: a [bsky, 1 points, 1 comments]
- Hey, if you get a minute, this paper on AI pre-deployment-tuning is fucking fire. arxiv.org/pdf/2501.19393 [bsky, 0 points, 0 comments]
- Interesting paper from Niklas Muennighoff, Zitong Yang et al. from Stanford arxiv.org/abs/2501.19393 "Thus, we ask: what is the simplest approach to achieve both test-time scaling and strong reasoni [bsky, 0 points, 0 comments]
- s1: Simple test-time scaling #llms #reasoning #wait #s1 #thinking [bsky, 0 points, 0 comments]
- S1, a 50 dollar o1 competitor arxiv.org/pdf/2501.19393 Looks like @profgalloway.com is right about the OpenAI bubble. [bsky, 0 points, 0 comments]
- Yet another LLM. s1 is developed by Washington and Stanford universities and it's free. arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- Adding "Wait" multiple times increased accuracy even further (see image). If you've had any experience using DeepSeek, then you've probably already seen this behaviour. This may explain why. Paper [bsky, 0 points, 0 comments]
- Stanford / UW Team showcases accuracy improvements in reasoning through test-time scaling, high quality data, and brute forcing reasoning backtracking on Qwen 32B fine tune arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
- Simple test-time scaling [Muennighoff+, 2025] The authors reproduced the test-time scaling curve by fine-tuning Qwen2.5-32B with s1K, a set of reasoning traces generated by Gemini. They controlled the [bsky, 0 points, 0 comments]
- [2501.19393] s1: Simple test-time scaling arxiv.org/abs/2501.19393 [bsky, 0 points, 0 comments]
Related