Blagden, Chase
- Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
2025/01/08 by Violet Xiang, Charlie Snell, Xiang, Violet +25 · 17 voices · 17 citations
#cs.AI #cs.CL
- Generative Reward Models
2024/10/02 by Mahan, Dakota, Van Phung, Duy, Rafailov, Rafael +6 · 29 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
2025/02/24 by Alon Albalak, Duy Phung, Albalak, Alon +18 · 28 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models
2024/12/04 by Alex Havrilla, Andrew M. Dai, Havrilla, Alex +36 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
2025/06/05 by Violet Xiang, Xiang, Violet, Rafael Rafailov +10 · 12 citations
Computer Science · #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Explainable Artificial Intelligence (XAI)