vix.ing · top · new · best · stats · spec

Rafael Rafailov

  1. Direct Preference Optimization: Your Language Model is Secretly a Reward Model
    2023/05/29 by Rafael Rafailov, Rafailov, Rafael, Archit Sharma +9 · 9 voices · 1515 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  2. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
    2024/04/01 by Matthias Gerstgrasser, Rylan Schaeffer, Gerstgrasser, Matthias +26 · 16 voices · 31 citations
    Computer Science · #Semantic Web and Ontologies #cs.AI #cs.CL #cs.ET #cs.LG #stat.ML
  3. Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
    2025/01/08 by Violet Xiang, Charlie Snell, Xiang, Violet +25 · 17 voices · 16 citations
    #cs.AI #cs.CL
  4. OpenVLA: An Open-Source Vision-Language-Action Model
    2024/06/13 by Moo Jin Kim, Kim, Moo Jin, Karl Pertsch +33 · 618 citations
    Computer Science · #Semantic Web and Ontologies
  5. Diffusion Model Alignment Using Direct Preference Optimization
    2023/11/21 by Bram Wallace, Wallace, Bram, Meihua Dang +17 · 161 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Recommender Systems and Techniques #Topic Modeling
  6. Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
    2023/05/24 by Katherine Tian, Tian, Katherine, Eric Mitchell +13 · 114 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling
  7. PERSONA: A Reproducible Testbed for Pluralistic Alignment
    2024/07/24 by Louis Castricato, Nathan Lile, Castricato, Louis +7 · 1 voice · 16 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
  8. COMBO: Conservative Offline Model-Based Policy Optimization
    2021/02/16 by Tianhe Yu, Aviral Kumar, Yu, Tianhe +9 · 23 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Adversarial Robustness in Machine Learning #Advanced Bandit Algorithms Research
  9. From r to Q^*: Your Language Model is Secretly a Q-Function
    2024/04/18 by Rafael Rafailov, Joey Hejna, Rafailov, Rafael +5 · 27 citations
    Computer Science · #Algorithms and Data Compression #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  10. Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
    2024/04/22 by Fahim Tajwar, Anikait Singh, Tajwar, Fahim +15 · 24 citations
    Decision Sciences · Economics, Econometrics and Finance · #Efficiency Analysis Using DEA #Healthcare Policy and Management #Auction Theory and Applications
  11. LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
    2025/07/01 by Daniel Fein, Fein, Daniel, Sebastian Russo +9 · 2 voices · 10 citations
    #cs.CL #cs.AI
  12. Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World
    2024/10/22 by Joshua Kazdan, Kazdan, Joshua, Rylan Schaeffer +11 · 2 voices · 9 citations
    #cs.LG #cs.AI
  13. Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
    2025/02/24 by Alon Albalak, Albalak, Alon, Duy Phung +18 · 28 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  14. Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
    2024/06/05 by Rafael Rafailov, Rafailov, Rafael, Yaswanth Chittepu +13 · 15 citations
    Engineering · Decision Sciences · Computer Science · #Advanced Control Systems Optimization #Simulation Techniques and Applications #Statistical and Computational Modeling
  15. An Emulator for Fine-Tuning Large Language Models using Small Language Models
    2023/10/19 by Eric Mitchell, Mitchell, Eric, Rafael Rafailov +7 · 8 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
  16. Vision-Based Manipulators Need to Also See from Their Hands
    2022/03/15 by Kyle Hsu, Hsu, Kyle, Moo Jin Kim +7 · 3 citations
    Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Robotics (cs.RO)
  17. Offline Meta-Reinforcement Learning with Advantage Weighting
    2020/08/13 by Eric Mitchell, Mitchell, Eric, Rafael Rafailov +7 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Reinforcement Learning in Robotics
  18. Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
    2025/06/05 by Violet Xiang, Xiang, Violet, Rafael Rafailov +10 · 12 citations
    Computer Science · #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Explainable Artificial Intelligence (XAI)
  19. Offline Regularised Reinforcement Learning for Large Language Models Alignment
    2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  20. Visual Adversarial Imitation Learning using Variational Models
    2021/07/16 by Rafael Rafailov, Rafailov, Rafael, Tianhe Yu +5 · 2 citations
    Computer Science · Physics and Astronomy · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics #Robotics (cs.RO)
  21. Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
    2024/04/22 by Jan-Philipp Fränken, Eric Zelikman, Fränken, Jan-Philipp +9 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  22. D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
    2024/08/15 by Rafael Rafailov, Rafailov, Rafael, Kyle Hatch +21 · 2 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO)