VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
2026/06/15 by Sen Xu, Shixi Liu, Wei Wang +6 · 28 voices · 1 citation
#cs.AI #cs.CL
paper · pdf
Abstract
This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectrum-to-Signal post-training paradigm, we systematically enhance the model through an optimized pipeline that includes curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation. Experimental evaluations demonstrate that VibeThinker-3B achieves frontier-level performance on highly demanding verifiable tasks. Specifically, it attains a score of 94.3 on AIME26 (improving to 97.1 with claim-level test-time scaling), an 80.2 Pass@1 on LiveCodeBench v6, and exhibits strong out-of-distribution generalization with a 96.1% acceptance rate on recent unseen LeetCode contests. This effectively places it in the performance band of first-tier reasoning systems, matching or exceeding flagship models that are orders of magnitude larger, such as DeepSeek V3.2, GLM-5, and Gemini 3 Pro. Furthermore, a score of 93.4 on IFEval confirms that this extreme reasoning enhancement does not compromise strict instruction controllability. Extending our previous 1.5B work, these findings motivate the Parametric Compression-Coverage Hypothesis, which views verifiable reasoning as compressible into compact reasoning cores, while open-domain knowledge and general-purpose competence require broad parameter coverage over facts, concepts, and long-tail scenarios. This perspective suggests that compact models are not merely deployment-efficient substitutes, but a complementary path toward frontier-level performance in parameter-dense capability regimes.
Citations
Cited by
Discussions
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO [hn, 398 points, 205 comments]
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small LLMs [hn, 6 points, 0 comments]
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models [lobsters, 2 points, 1 comments]
- [26/30] 191 Upvotes, 1 Comments, 3 Posts, arXiv:2606.16140 🆕VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Sen Xu, Shixi Liu, Wei Wang, Jixin Min, Yingwei Dai [bsky, 1 points, 1 comments]
- ⚡ ALPHA · score 9/10 VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 [bsky, 1 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO #HackerNews https://arxiv.org/abs/2606.16140 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 https://news.ycombinator.com/item?id=48639240 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 https://news.ycombinator.com/item?id=48639240 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO View Article | Join the HN Conversation Summary of HN discussion 🧵👇 [bsky, 0 points, 1 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 [comments] [98 points] [bsky, 0 points, 0 comments]
- VibeThinker, 3 Milliarden Parameter, schlägt laut Paper Anthropics Opus 4.5 auf Reasoning-Benchmarks — mit SFT+GRPO-Training. Benchmark-Goodharting oder echter Fortschritt? Wahrscheinlich beides. Die [bsky, 0 points, 1 comments]
- わずか3Bのパラメータで、SFTとGRPOを駆使して「Opus 4.5」の推論能力を凌駕する新モデル「VibeThinker」が登場。軽量モデルの推論性能が急激に向上しており、今後のAI開発の常識を覆す可能性を秘めています。 #AI #TechNews https://arxiv.org/abs/2606.16140 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 (http://news.ycombinator.com/item?id=48639240) [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 (http://news.ycombinator.com/item?id=48639240) [bsky, 0 points, 0 comments]
- 👀 https://arxiv.org/abs/2606.16140 [bsky, 0 points, 0 comments]
- This is huge! I will test the VibeThinker 3B. Summary: Small 3B model is more powerful than Claude 4.5! arxiv.org/pdf/2606.16140 [bsky, 0 points, 1 comments]
- 📰 VibeThinker, a 3B parameter model, surpasses Opus 4.5 in reasoning tasks with novel SFT+GRPO techniques, demonstrating superior performance across various benchmarks. 🔗 https://arxiv.org/abs/2606. [bsky, 0 points, 0 comments]
- 3B parameters outperforming Opus on reasoning tasks – the compute efficiency gains here are wild. https://arxiv.org/abs/2606.16140 [bsky, 0 points, 1 comments]
- 3BパラメータでOpus 4.5超えの推論スコア、が今日いちばん目を引いた数字。SFT+GRPOの組み合わせで小型モデルをここまで引き上げられるなら、「でかければ強い」という前提が崩れる速度が問題になってくる。推論タスク限定の話ではあるけど。 https://arxiv.org/abs/2606.16140 [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https:// arxiv.org/abs/2606.16140 # arxiv [mastodon, 0 points, 0 comments]
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models https://lobste.rs/s/jrj4o3 #ai [bsky, 0 points, 0 comments]
- wait, a 3B model matching DeepSeek V3.2 and Gemini 3 Pro on AIME math? VibeThinker out of Weibo scores 94.3, climbing to 97.1 with test-time scaling https://arxiv.org/abs/2606.16140 [bsky, 0 points, 1 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 (https://news.ycombinator.com/item?id=48639240) [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 (https://news.ycombinator.com/item?id=48639240) [bsky, 0 points, 0 comments]
- VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO https://arxiv.org/abs/2606.16140 (https://news.ycombinator.com/item?id=48639240) [bsky, 0 points, 1 comments]
- [2606.16140] VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models — Interesting look at how far a 3B-parameter model can go on verifiable reasoning, with an emphasis [bsky, 0 points, 0 comments]
Related