vix.ing · top · new · best · stats · spec

Reducing Variance Caused by Communication in Decentralized Multi-agent Deep Reinforcement Learning

2025/02/10 by Changxi Zhu, Zhu, Changxi, Mehdi Dastani +3 · 1 citation
Computer Science · Engineering · #68T05 #Advanced Research in Systems and Signal Processing #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2502.06261

openalex publication_date 2025/02/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In decentralized multi-agent deep reinforcement learning (MADRL), communication can help agents to gain a better understanding of the environment to better coordinate their behaviors. Nevertheless, communication may involve uncertainty, which potentially introduces variance to the learning of decentralized agents. In this paper, we focus on a specific decentralized MADRL setting with communication and conduct a theoretical analysis to study the variance that is caused by communication in policy gradients. We propose modular techniques to reduce the variance in policy gradients during training. We adopt our modular techniques into two existing algorithms for decentralized MADRL with communication and evaluate them on multiple tasks in the StarCraft Multi-Agent Challenge and Traffic Junction domains. The results show that decentralized MADRL communication methods extended with our proposed techniques not only achieve high-performing agents but also reduce variance in policy gradients during training.

Cited by

Related