vix.ing · top · new · best · stats · spec

Weixun Wang

  1. Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem
    2025/12/31 by Weixun Wang, XiaoXiao Xu, Wanhe An +86 · 22 voices · 2 citations
    #cs.AI #cs.CL
  2. Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
    2025/08/11 by Zihe Liu, Jiashun Liu, Liu, Zihe +29 · 2 voices · 22 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.CL #cs.LG
  3. Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping
    2020/11/05 by Yujing Hu, Weixun Wang, Hu, Yujing +13 · 12 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Software Engineering Research
  4. Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
    2025/06/06 by Weixun Wang, Wang, Weixun, Xiong, Shaopan +73 · 36 citations
    Computer Science · #Reinforcement Learning in Robotics #Machine Learning and Data Classification #Stochastic Gradient Optimization Techniques
  5. A2C is a special case of PPO
    2022/05/18 by Huang, Shengyi, Anssi Kanervisto, Antonin Raffin +8 · 3 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Artificial Intelligence in Games #Advanced Bandit Algorithms Research
  6. Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
    2025/03/20 by Luo, Yijia, Song, Yulin, Xingyao Zhang +10 · 10 citations
    Engineering · #Advanced Control Systems Optimization #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Process Optimization and Integration
  7. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
    2025/02/23 by Alexander Zhang, Zhang, Alexander, Dong, Marcus +32 · 4 citations
    Computer Science · #Software Engineering Research #Software System Performance and Reliability #Topic Modeling
  8. Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
    2025/05/26 by Boren Zheng, Zheng, Baihui, Zheng, Boren +19 · 5 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Semantic Web and Ontologies
  9. RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
    2025/09/25 by Wei Gao, Gao, Wei, Yuheng Zhao +24 · 7 citations
    Computer Science · Engineering · #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Real-Time Systems Scheduling #Real-time simulation and control systems #Software Reliability and Analysis Research #and Cluster Computing (cs.DC)
  10. Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
    2025/10/02 by Liu, Jiashun, Johan Obando-Ceron, Lu Han +16 · 2 citations
    Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG)
  11. ProgCo: Program Helps Self-Correction of Large Language Models
    2025/01/02 by Xiaoshuai Song, Song, Xiaoshuai, Yanan Wu +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling