Zhu, Zixu
- UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function
2024/10/28 by Wang, Zhichao, Bi, Bin, Zhu, Zixu +3 · 4 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)