vix.ing · top · new · best · stats · spec

Pierre Harvey Richemond

  1. Generalized Preference Optimization: A Unified Approach to Offline Alignment
    2024/02/08 by Yunhao Tang, Tang, Yunhao, Zhaohan Daniel Guo +17 · 19 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Constraint Satisfaction and Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Criteria Decision Making
  2. Scaling Instructable Agents Across Many Simulated Worlds
    2024/03/13 by SIMA team, Maria Abi Raad, SIMA Team +184 · 16 citations
    Computer Science · #Robotic Path Planning Algorithms #Teaching and Learning Programming
  3. Human Alignment of Large Language Models through Online Preference Optimisation
    2024/03/13 by Daniele Calandriello, Calandriello, Daniele, Daniel Guo +23 · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  4. Offline Regularised Reinforcement Learning for Large Language Models Alignment
    2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling