vix.ing · top · new · best · stats · spec

Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field\n Control/Game in Continuous Time

2020/08/16 by Wei-Chen Wang, Wang, Weichen, Jiequn Han +6
Computer Science · #Reinforcement Learning in Robotics #Adaptive Dynamic Programming Control #Machine Learning and ELM

paper · pdf · doi:10.48550/arxiv.2008.06845

Abstract

Reinforcement learning is a powerful tool to learn the optimal policy of\npossibly multiple agents by interacting with the environment. As the number of\nagents grow to be very large, the system can be approximated by a mean-field\nproblem. Therefore, it has motivated new research directions for mean-field\ncontrol (MFC) and mean-field game (MFG). In this paper, we study the policy\ngradient method for the linear-quadratic mean-field control and game, where we\nassume each agent has identical linear state transitions and quadratic cost\nfunctions. While most of the recent works on policy gradient for MFC and MFG\nare based on discrete-time models, we focus on the continuous-time models where\nsome analyzing techniques can be interesting to the readers. For both MFC and\nMFG, we provide policy gradient update and show that it converges to the\noptimal solution at a linear rate, which is verified by a synthetic simulation.\nFor MFG, we also provide sufficient conditions for the existence and uniqueness\nof the Nash equilibrium.\n

Citations

Related