Spurious Rewards: Rethinking Training Signals in RLVR
2025/06/12 by Rulin Shao, Shuyue Stella Li, Shao, Rulin +26 · 2 voices · 91 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2506.10947
openalex publication_date 2025/06/12 · arxiv published 2025/06/12 · openalex created_date 2025/10/10 · arxiv updated 2026/02/25 · openalex updated_date 2026/07/28
Abstract
We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or even negative correlation with the correct answer. For example, RLVR training with GRPO improves MATH-500 performance for Qwen2.5-Math-7B by 21.4 percentage points using randomly assigned rewards, nearly matching the 29.1-point gain from ground-truth rewards. To explain this counterintuitive observation, we show that GRPO exhibits a clipping bias from the clip term, which can amplify high-prior behaviors learned during pretraining even without informative rewards. As a case study, we identify one such behavior in Qwen2.5-Math models, which we call code reasoning -- reasoning in code without actual code execution; code-reasoning frequency increases from 65 percent to over 90 percent with spurious rewards. However, the presence of such amplifiable behaviors is highly model-dependent. In practice, spurious rewards that are effective for Qwen models often fail to produce gains for other model families, such as Llama3 or OLMo2. Our results highlight the importance of validating RL methods across diverse models rather than relying on a single de facto choice: large gains can arise on Qwen models even from random rewards that do not reflect genuine capability improvements.
Citations
Cited by
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
- What Is Preference Optimization Doing, How and Why?
- ThetaEvolve: Test-time Learning on Open Problems
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- GPS: General Per-Sample Prompter
- GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards
- Better LLM Reasoning via Dual-Play
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
- Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
- Limits of Generalization in RLVR: Two Case Studies in Mathematical Reasoning
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- VAR: Visual Attention Reasoning via Structured Search and Backtracking
- Rethinking On-policy Optimization for Query Augmentation
- LSPO: Length-aware Dynamic Sampling for Policy Optimization in LLM Reasoning
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Can GRPO Help LLMs Transcend Their Pretraining Origin?
- MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
- SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
- Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL
- GCPO: When Contrast Fails, Go Gold
- Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
- AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning
- Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
- The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- Nudging the Boundaries of LLM Reasoning
- Humanline: Online Alignment as Perceptual Loss
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasks
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- Noisy Data is Destructive to Reinforcement Learning with Verifiable Rewards
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- The Need for Verification in AI-Driven Scientific Discovery
- Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
- Self-Rewarding Vision-Language Model via Reasoning Decomposition
- ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
- Self-Questioning Language Models
- QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
- Post-Completion Learning for Language Models
- GUI-G2: Gaussian Reward Modeling for GUI Grounding
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
- Resa: Transparent Reasoning Models via SAEs
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
- The Serial Scaling Hypothesis
- KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
- First Return, Entropy-Eliciting Explore
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
- ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation
- OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
- ReDit: Reward Dithering for Improved LLM Policy Optimization
- HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges
- Play to Generalize: Learning to Reason Through Game Play
- Reasoning with Exploration: An Entropy Perspective
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
- Improving LLM-Generated Code Quality with GRPO
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- Maximizing Confidence Alone Improves Reasoning
- Can Large Reasoning Models Self-Train?
- Learning to Reason without External Rewards
- Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
- Steering LLM Reasoning Through Bias-Only Adaptation
- Towards Revealing the Effectiveness of Small-Scale Fine-tuning in R1-style Reinforcement Learning
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- PRL: Prompts from Reinforcement Learning
- Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
- Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
- RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
- Reasoning with Sampling: Cutting at Decision Points
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning
- Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
- Reward Hacking in Rubric-Based Reinforcement Learning
- SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
- An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
- CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
- TTRL: Test-Time Reinforcement Learning
- HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Discussions
Related