Weixun Wang
- Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
2025/12/31 by Weixun Wang, XiaoXiao Xu, Wanhe An +86 · 22 voices · 2 citations
#cs.AI #cs.CL
- Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
2025/08/11 by Zihe Liu, Jiashun Liu, Liu, Zihe +29 · 2 voices · 22 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.CL #cs.LG
- Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping
2020/11/05 by Yujing Hu, Weixun Wang, Hu, Yujing +13 · 12 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Software Engineering Research
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
2025/06/06 by Weixun Wang, Wang, Weixun, Xiong, Shaopan +73 · 36 citations
Computer Science · #Reinforcement Learning in Robotics #Machine Learning and Data Classification #Stochastic Gradient Optimization Techniques
- A2C is a special case of PPO
2022/05/18 by Huang, Shengyi, Anssi Kanervisto, Antonin Raffin +8 · 3 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Artificial Intelligence in Games #Advanced Bandit Algorithms Research
- Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
2025/03/20 by Luo, Yijia, Song, Yulin, Xingyao Zhang +10 · 10 citations
Engineering · #Advanced Control Systems Optimization #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Process Optimization and Integration
- CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
2025/02/23 by Alexander Zhang, Zhang, Alexander, Dong, Marcus +32 · 4 citations
Computer Science · #Software Engineering Research #Software System Performance and Reliability #Topic Modeling
- Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
2025/05/26 by Boren Zheng, Zheng, Baihui, Zheng, Boren +19 · 5 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Semantic Web and Ontologies
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
2025/09/25 by Wei Gao, Gao, Wei, Yuheng Zhao +24 · 7 citations
Computer Science · Engineering · #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Real-Time Systems Scheduling #Real-time simulation and control systems #Software Reliability and Analysis Research #and Cluster Computing (cs.DC)
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
2025/10/02 by Liu, Jiashun, Johan Obando-Ceron, Lu Han +16 · 2 citations
Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG)
- ProgCo: Program Helps Self-Correction of Large Language Models
2025/01/02 by Xiaoshuai Song, Song, Xiaoshuai, Yanan Wu +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling