vix.ing · top · new · best · stats · spec

Stochastic Recursive Momentum for Policy Gradient Methods

2020/03/09 by Huizhuo Yuan, Yuan, Huizhuo, Xiangru Lian +5 · 3 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Search Problems #Reinforcement Learning in Robotics #Stochastic Gradient Optimization Techniques

paper · pdf · doi:10.48550/arxiv.2003.04302

openalex publication_date 2020/03/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced policy gradient in an exponential moving average fashion. STORM-PG enjoys a provably sharp O(1/ε3) sample complexity bound for STORM-PG, matching the best-known convergence rate for policy gradient algorithm. In the mean time, STORM-PG avoids the alternations between large batches and small batches which persists in comparable variance-reduced policy gradient methods, allowing considerably simpler parameter tuning. Numerical experiments depicts the superiority of our algorithm over comparative policy gradient algorithms.

Citations

Cited by

Related