vix.ing · top · new · best · stats

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

2025/05/28 by Youssef Mroueh, Mroueh, Youssef, Nicolas Dupuis +15 · 11 citations
Psychology · #FOS: Computer and information sciences #Function (biology) #Group (periodic table) #Human Resource Development and Performance Evaluation #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement learning #Sample (material) #Sampling (signal processing) #Training (meteorology) #Verifiable secret sharing

paper · pdf · doi:10.48550/arxiv.2505.22257

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/05/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Optimization (PPO), which improves training stability, sampling efficiency, and memory usage. In addition, a recent analysis of GRPO suggests that estimating the advantage function with off-policy samples could be beneficial. Building on these observations, we adapt GRPO to the off-policy setting. We show that both on-policy and off-policy GRPO objectives yield an improvement in the reward. This result motivates the use of clipped surrogate objectives in the off-policy version of GRPO. We then compare the empirical performance of reinforcement learning with verifiable rewards in post-training using both GRPO variants. Our results show that off-policy GRPO either significantly outperforms or performs on par with its on-policy counterpart.

Citations

Cited by

Related