2021/09/18 by Chapman Siu, Siu, Chapman, Jason Traish +3
Computer Science · #Adaptive Dynamic Programming Control #Distributed Control Multi-Agent Systems #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mobile Crowdsensing and Crowdsourcing #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2109.09038
openalex publication_date 2021/09/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We propose using regularization for Multi-Agent Reinforcement Learning rather\nthan learning explicit cooperative structures called em Multi-Agent\nRegularized Q-learning (MARQ). Many MARL approaches leverage centralized\nstructures in order to exploit global state information or removing\ncommunication constraints when the agents act in a decentralized manner.\nInstead of learning redundant structures which is removed during agent\nexecution, we propose instead to leverage shared experiences of the agents to\nregularize the individual policies in order to promote structured exploration.\nWe examine several different approaches to how MARQ can either explicitly or\nimplicitly regularize our policies in a multi-agent setting. MARQ aims to\naddress these limitations in the MARL context through applying regularization\nconstraints which can correct bias in off-policy out-of-distribution agent\nexperiences and promote diverse exploration. Our algorithm is evaluated on\nseveral benchmark multi-agent environments and we show that MARQ consistently\noutperforms several baselines and state-of-the-art algorithms; learning in\nfewer steps and converging to higher returns.\n