Rafael Rafailov
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
2023/05/29 by Rafael Rafailov, Rafailov, Rafael, Archit Sharma +9 · 9 voices · 1515 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
2024/04/01 by Matthias Gerstgrasser, Rylan Schaeffer, Gerstgrasser, Matthias +26 · 16 voices · 31 citations
Computer Science · #Semantic Web and Ontologies #cs.AI #cs.CL #cs.ET #cs.LG #stat.ML
- Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
2025/01/08 by Violet Xiang, Charlie Snell, Xiang, Violet +25 · 17 voices · 16 citations
#cs.AI #cs.CL
- OpenVLA: An Open-Source Vision-Language-Action Model
2024/06/13 by Moo Jin Kim, Kim, Moo Jin, Karl Pertsch +33 · 618 citations
Computer Science · #Semantic Web and Ontologies
- Diffusion Model Alignment Using Direct Preference Optimization
2023/11/21 by Bram Wallace, Wallace, Bram, Meihua Dang +17 · 161 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Recommender Systems and Techniques #Topic Modeling
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
2023/05/24 by Katherine Tian, Tian, Katherine, Eric Mitchell +13 · 114 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling
- PERSONA: A Reproducible Testbed for Pluralistic Alignment
2024/07/24 by Louis Castricato, Nathan Lile, Castricato, Louis +7 · 1 voice · 16 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
- COMBO: Conservative Offline Model-Based Policy Optimization
2021/02/16 by Tianhe Yu, Aviral Kumar, Yu, Tianhe +9 · 23 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Adversarial Robustness in Machine Learning #Advanced Bandit Algorithms Research
- From r to Q^*: Your Language Model is Secretly a Q-Function
2024/04/18 by Rafael Rafailov, Joey Hejna, Rafailov, Rafael +5 · 27 citations
Computer Science · #Algorithms and Data Compression #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
2024/04/22 by Fahim Tajwar, Anikait Singh, Tajwar, Fahim +15 · 24 citations
Decision Sciences · Economics, Econometrics and Finance · #Efficiency Analysis Using DEA #Healthcare Policy and Management #Auction Theory and Applications
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
2025/07/01 by Daniel Fein, Fein, Daniel, Sebastian Russo +9 · 2 voices · 10 citations
#cs.CL #cs.AI
- Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
2024/10/22 by Joshua Kazdan, Kazdan, Joshua, Rylan Schaeffer +11 · 2 voices · 9 citations
#cs.LG #cs.AI
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
2025/02/24 by Alon Albalak, Albalak, Alon, Duy Phung +18 · 28 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
2024/06/05 by Rafael Rafailov, Rafailov, Rafael, Yaswanth Chittepu +13 · 15 citations
Engineering · Decision Sciences · Computer Science · #Advanced Control Systems Optimization #Simulation Techniques and Applications #Statistical and Computational Modeling
- An Emulator for Fine-Tuning Large Language Models using Small Language Models
2023/10/19 by Eric Mitchell, Mitchell, Eric, Rafael Rafailov +7 · 8 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- Vision-Based Manipulators Need to Also See from Their Hands
2022/03/15 by Kyle Hsu, Hsu, Kyle, Moo Jin Kim +7 · 3 citations
Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Robotics (cs.RO)
- Offline Meta-Reinforcement Learning with Advantage Weighting
2020/08/13 by Eric Mitchell, Mitchell, Eric, Rafael Rafailov +7 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Reinforcement Learning in Robotics
- Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
2025/06/05 by Violet Xiang, Xiang, Violet, Rafael Rafailov +10 · 12 citations
Computer Science · #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Explainable Artificial Intelligence (XAI)
- Offline Regularised Reinforcement Learning for Large Language Models Alignment
2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Visual Adversarial Imitation Learning using Variational Models
2021/07/16 by Rafael Rafailov, Rafailov, Rafael, Tianhe Yu +5 · 2 citations
Computer Science · Physics and Astronomy · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics #Robotics (cs.RO)
- Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
2024/04/22 by Jan-Philipp Fränken, Eric Zelikman, Fränken, Jan-Philipp +9 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
2024/08/15 by Rafael Rafailov, Rafailov, Rafael, Kyle Hatch +21 · 2 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO)