vix.ing · top · new · best · stats · spec

Zhu, Zixu

  1. UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function
    2024/10/28 by Wang, Zhichao, Bi, Bin, Zhu, Zixu +3 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)