2022/09/06 by Jinchi Chen, Chen, Jinchi, Jie Feng +5 · 2 citations
Computer Science · Engineering · #Advanced MIMO Systems Optimization #Distributed Control Multi-Agent Systems #FOS: Mathematics #Optimization and Control (math.OC)
paper · pdf · doi:10.48550/arxiv.2209.02179
openalex publication_date 2022/09/06 · openalex created_date 2022/09/08 · openalex updated_date 2026/07/28
This paper studies a policy optimization problem arising from collaborative multi-agent reinforcement learning in a decentralized setting where agents communicate with their neighbors over an undirected graph to maximize the sum of their cumulative rewards. A novel decentralized natural policy gradient method, dubbed Momentum-based Decentralized Natural Policy Gradient (MDNPG), is proposed, which incorporates natural gradient, momentum-based variance reduction, and gradient tracking into the decentralized stochastic gradient ascent framework. The O(n-1ε-3) sample complexity for MDNPG to converge to an ε-stationary point has been established under standard assumptions, where n is the number of agents. It indicates that MDNPG can achieve the optimal convergence rate for decentralized policy gradient methods and possesses a linear speedup in contrast to centralized optimization methods. Moreover, superior empirical performance of MDNPG over other state-of-the-art algorithms has been demonstrated by extensive numerical experiments.