2022/02/18 by Dingyang Chen, Yile Li, Chen, Dingyang +3 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics #cs.AI #cs.GT #cs.LG #cs.MA
paper · pdf · doi:10.48550/arxiv.2202.09422
openalex publication_date 2022/02/18 · arxiv created 2022/03/31 · arxiv updated 2022/04/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recent success in cooperative multi-agent reinforcement learning (MARL) relies on centralized training and policy sharing. Centralized training eliminates the issue of non-stationarity MARL yet induces large communication costs, and policy sharing is empirically crucial to efficient learning in certain tasks yet lacks theoretical justification. In this paper, we formally characterize a subclass of cooperative Markov games where agents exhibit a certain form of homogeneity such that policy sharing provably incurs no suboptimality. This enables us to develop the first consensus-based decentralized actor-critic method where the consensus update is applied to both the actors and the critics while ensuring convergence. We also develop practical algorithms based on our decentralized actor-critic method to reduce the communication cost during training, while still yielding policies comparable with centralized training.