vix.ing · top · new · best · stats · spec

Zhou, Runlong

  1. Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments
    2023/01/31 by Zhou, Runlong, Zhang, Zihan, Du, Simon S. · 4 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  2. Stochastic Shortest Path: Minimax, Parameter-Free and Towards Horizon-Free Regret
    2021/04/22 by Tarbouriech, Jean, Zhou, Runlong, Du, Simon S. +3 · 2 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Preference-Based Multi-Agent Reinforcement Learning: Data Coverage and Algorithmic Techniques
    2024/09/01 by N Zhang, Xinqi Wang, Zhang, Natalia +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics
  4. Transformers are Efficient Compilers, Provably
    2024/10/07 by Xiyu Zhai, Zhai, Xiyu, Runlong Zhou +5 · 1 voice
    #cs.PL #cs.LG
  5. Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
    2024/02/20 by Runlong Zhou, Zhou, Runlong, Simon S. Du +3 · 2 citations
    Engineering · #Computation and Language (cs.CL) #Experimental Learning in Engineering #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. The Crucial Role of Samplers in Online Direct Preference Optimization
    2024/09/29 by Shi, Ruizhe, Zhou, Runlong, Du, Simon S. · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. Horizon-Free and Variance-Dependent Reinforcement Learning for Latent Markov Decision Processes
    2022/10/20 by Zhou, Runlong, Wang, Ruosong, Du, Simon S. · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  8. Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised Learning
    2023/10/30 by Zhou, Zhaoyi, Zhu, Chuning, Zhou, Runlong +3 · 1 citation
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
    2025/03/11 by Zhou, Runlong, Fazel, Maryam, Du, Simon S. · 2 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  10. Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
    2025/05/26 by Ruizhe Shi, Shi, Ruizhe, Runlong Zhou +8 · 2 citations
    Decision Sciences · Computer Science · #Multi-Criteria Decision Making #Semantic Web and Ontologies