vix.ing · top · new · best · stats · spec

Yibo Xie

  1. Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning
    2026/08/02 by Wenhao Zhang, Yibo Xie, Rui Wang +9
    Computer Science · #cs.AI #cs.LG