vix.ing · top · new · best · stats · spec

Shamir, Gil

  1. Offline Regularised Reinforcement Learning for Large Language Models Alignment
    2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 12 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling