Yingxiang Yang
- Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
2024/05/26 by Zhihan Liu, Liu, Zhihan, Miao Lu +13 · 33 citations
Computer Science · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Digital Filter Design and Implementation #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Numerical Methods and Algorithms
- Let Models Speak Ciphers: Multiagent Debate through Embeddings
2023/10/10 by Chau Pham, Boyi Liu, Pham, Chau +15 · 17 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Reason out Your Layout: Evoking the Layout Master from Large Language Models for Text-to-Image Synthesis
2023/11/28 by Yongfei Liu, Chen, Xiaohui, Liu, Yongfei +10 · 4 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
- How Can LLM Guide RL? A Value-Based Approach
2024/02/25 by Shenao Zhang, Sirui Zheng, Zhang, Shenao +15 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Digital Rights Management and Security #FOS: Computer and information sciences #Library Science and Information Systems #Machine Learning (cs.LG) #Mathematics, Computing, and Information Processing
- Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
2024/10/10 by Shenao Zhang, Zhang, Shenao, Zhihan Liu +15 · 3 citations
Computer Science · #Digital Rights Management and Security