Zhou, Runlong
- Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments
2023/01/31 by Zhou, Runlong, Zhang, Zihan, Du, Simon S. · 4 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Stochastic Shortest Path: Minimax, Parameter-Free and Towards Horizon-Free Regret
2021/04/22 by Tarbouriech, Jean, Zhou, Runlong, Du, Simon S. +3 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Preference-Based Multi-Agent Reinforcement Learning: Data Coverage and Algorithmic Techniques
2024/09/01 by N Zhang, Xinqi Wang, Zhang, Natalia +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics
- Transformers are Efficient Compilers, Provably
2024/10/07 by Xiyu Zhai, Zhai, Xiyu, Runlong Zhou +5 · 1 voice
#cs.PL #cs.LG
- Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
2024/02/20 by Runlong Zhou, Zhou, Runlong, Simon S. Du +3 · 2 citations
Engineering · #Computation and Language (cs.CL) #Experimental Learning in Engineering #FOS: Computer and information sciences #Machine Learning (cs.LG)
- The Crucial Role of Samplers in Online Direct Preference Optimization
2024/09/29 by Shi, Ruizhe, Zhou, Runlong, Du, Simon S. · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Horizon-Free and Variance-Dependent Reinforcement Learning for Latent Markov Decision Processes
2022/10/20 by Zhou, Runlong, Wang, Ruosong, Du, Simon S. · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning
2023/10/30 by Zhou, Zhaoyi, Zhu, Chuning, Zhou, Runlong +3 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
2025/03/11 by Zhou, Runlong, Fazel, Maryam, Du, Simon S. · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
2025/05/26 by Ruizhe Shi, Shi, Ruizhe, Runlong Zhou +8 · 2 citations
Decision Sciences · Computer Science · #Multi-Criteria Decision Making #Semantic Web and Ontologies