Tan, Hongze
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
2025/08/06 by Tan, Hongze, Pan, Jianfei, Lin, Jinghao +4 · 12 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences