vix.ing · top · new · best · stats · spec

On the Theory of Policy Gradient Methods: Optimality, Approximation, and\n Distribution Shift

2019/08/01 by Alekh Agarwal, Agarwal, Alekh, Sham M. Kakade +5 · 7 citations
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Fuel Cells and Related Materials #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.1908.00261

openalex publication_date 2019/08/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Policy gradient methods are among the most effective methods in challenging\nreinforcement learning problems with large state and/or action spaces. However,\nlittle is known about even their most basic theoretical convergence properties,\nincluding: if and how fast they converge to a globally optimal solution or how\nthey cope with approximation error due to using a restricted class of\nparametric policies. This work provides provable characterizations of the\ncomputational, approximation, and sample size properties of policy gradient\nmethods in the context of discounted Markov Decision Processes (MDPs). We focus\non both: "tabular" policy parameterizations, where the optimal policy is\ncontained in the class and where we show global convergence to the optimal\npolicy; and parametric policy classes (considering both log-linear and neural\npolicy classes), which may not contain the optimal policy and where we provide\nagnostic learning results. One central contribution of this work is in\nproviding approximation guarantees that are average case -- which avoid\nexplicit worst-case dependencies on the size of state space -- by making a\nformal connection to supervised learning under distribution shift. This\ncharacterization shows an important interplay between estimation error,\napproximation error, and exploration (as characterized through a precisely\ndefined condition number).\n

Citations

Cited by

Related