vix.ing · top · new · best · stats · spec

Zhang, Zizhuo

  1. Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
    2025/08/01 by Zizhuo Zhang, Zhang, Zizhuo, Jianing Zhu +15 · 9 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Natural Language Processing Techniques
  2. Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
    2025/05/15 by Zhang, Jiazheng, Jing, Wenqing, Zhang, Zizhuo +9 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)