Absolute Zero: Reinforced Self-play Reasoning with Zero Data
2025/05/06 by Andrew Zhao, Yiran Wu, Yilong Wu +20 · 23 voices · 76 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2505.03335
openalex publication_date 2025/05/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Reinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capabilities of large language models by learning directly from outcome-based rewards. Recent RLVR works that operate under the zero setting avoid supervision in labeling the reasoning process, but still depend on manually curated collections of questions and answers for training. The scarcity of high-quality, human-produced examples raises concerns about the long-term scalability of relying on human supervision, a challenge already evident in the domain of language model pretraining. Furthermore, in a hypothetical future where AI surpasses human intelligence, tasks provided by humans may offer limited learning potential for a superintelligent system. To address these concerns, we propose a new RLVR paradigm called Absolute Zero, in which a single model learns to propose tasks that maximize its own learning progress and improves reasoning by solving them, without relying on any external data. Under this paradigm, we introduce the Absolute Zero Reasoner (AZR), a system that self-evolves its training curriculum and reasoning ability by using a code executor to both validate proposed code reasoning tasks and verify answers, serving as an unified source of verifiable reward to guide open-ended yet grounded learning. Despite being trained entirely without external data, AZR achieves overall SOTA performance on coding and mathematical reasoning tasks, outperforming existing zero-setting models that rely on tens of thousands of in-domain human-curated examples. Furthermore, we demonstrate that AZR can be effectively applied across different model scales and is compatible with various model classes.
Citations
Cited by
Discussions
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [hn, 88 points, 19 comments]
- seems interesting, perhaps even promising arxiv.org/pdf/2505.03335 [bsky, 20 points, 2 comments]
- hahahahah we're so fucked arxiv.org/pdf/2505.03335 > This underscores the need for a new paradigm that begins to explore possibilities beyond the constraints of human-designed tasks and prepares for a [bsky, 4 points, 1 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [hn, 3 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [hn, 3 points, 0 comments]
- I had Claude compare your post to the Absolute Zero paper from yesterday and its conclusion was “Blog Post as Prophecy”… arxiv.org/pdf/2505.03335 [bsky, 3 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [hn, 3 points, 2 comments]
- Real Vernor Vinge vibes from this paper that learns to code from scratch via RL "Furthermore, in a hypothetical future where AI surpasses human intelligence, tasks provided by humans may offer limited [bsky, 2 points, 0 comments]
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data https://arxiv.org/abs/2505.03335v2 #AI #SelfPlayReasoning [bsky, 2 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [hn, 2 points, 0 comments]
- oi camarada eu venho aqui fazer um pedido atrevido pra vc (sigo o canal faz um tempo), tem chance de vc cobrir isso aqui? 🙏 arxiv.org/pdf/2505.03335 [bsky, 1 points, 0 comments]
- ⚡ Hackernews Top story: Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [bsky, 0 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data https://arxiv.org/abs/2505.03335 (https://news.ycombinator.com/item?id=43951985) [bsky, 0 points, 0 comments]
- Absolute Zero: the one where the model evolves through self-play without relying on external data. arxiv.org/abs/2505.03335 [bsky, 0 points, 1 comments]
- AI so smart it invents logic from nothing—too bad it’s arguing with itself about whether 2+2 is 5. www.arxiv.org/abs/25... [bsky, 0 points, 0 comments]
- AbsoluteZero: Reinforced Self-play Reasoningwith Zero Data arxiv.org/pdf/2505.03335 [bsky, 0 points, 0 comments]
- 今日読んだ論文 Absolute Zero: Reinforced Self-play Reasoning with Zero Data arxiv.org/abs/2505.03335 コーディング分野で、LLMが自分でタスク提案して自分で問いて強化学習をしていくという手法。コンセプト時点では抽象的に書かれているが、具体化されていくについて、やはりコーディング分野特有の性質(インタープリターで提 [bsky, 0 points, 0 comments]
- AbsoluteZero: Reinforced Self-play Reasoningwith Zero Data https://arxiv.org/pdf/2505.03335 [bsky, 0 points, 0 comments]
- Interesante paper: Este trabajo introduce un nuevo enfoque, llamado Absolute Zero, para entrenar modelos de razonamiento sin necesidad de datos etiquetados por humanos. arxiv.org/abs/2505.03335 [bsky, 0 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data https://arxiv.org/abs/2505.03335 (https://news.ycombinator.com/item?id=43951985) [bsky, 0 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data [bsky, 0 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data https://arxiv.org/abs/2505.03335 [bsky, 0 points, 0 comments]
- Absolute Zero: Reinforced Self-Play Reasoning with Zero Data https://arxiv.org/abs/2505.03335 [bsky, 0 points, 0 comments]
Related