2022/02/28 by Jing Dong, Dong, Jing, Li Shen +5
Computer Science · Mathematics · #Adaptive Dynamic Programming Control #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical Biology Tumor Growth #Reinforcement Learning in Robotics #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2202.13863
arxiv created 2022/02/28 · openalex publication_date 2022/02/28 · arxiv updated 2022/03/01 · openalex created_date 2022/04/03 · openalex updated_date 2026/07/28
We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascent is applied with an adaptive proximal term for robust learning rates. We show the first efficient convergence result with primal-dual actor-critic with a convergence rate of O(√((ln (N d G2 ))/(N))) under Markovian sampling, where G is the element-wise maximum of the gradient, N is the number of iterations, and d is the dimension of the gradient. Our result is presented with only the Polyak-Łojasiewicz condition for the dual variables, which is easy to verify and applicable to a wide range of reinforcement learning (RL) scenarios. The algorithm and analysis are general enough to be applied to other RL settings, like multi-agent RL. Empirical results on OpenAI Gym continuous control tasks corroborate our theoretical findings.