vix.ing · top · new · best · stats · spec

QTRAN: Learning to Factorize with Transformation for Cooperative\n Multi-Agent Reinforcement Learning

2019/05/14 by Kyunghwan Son, Son, Kyunghwan, Dae Woo Kim +7 · 61 citations
Computer Science · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms

paper · pdf · doi:10.48550/arxiv.1905.05408

openalex publication_date 2019/05/14 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28

Abstract

We explore value-based solutions for multi-agent reinforcement learning\n(MARL) tasks in the centralized training with decentralized execution (CTDE)\nregime popularized recently. However, VDN and QMIX are representative examples\nthat use the idea of factorization of the joint action-value function into\nindividual ones for decentralized execution. VDN and QMIX address only a\nfraction of factorizable MARL tasks due to their structural constraint in\nfactorization such as additivity and monotonicity. In this paper, we propose a\nnew factorization method for MARL, QTRAN, which is free from such structural\nconstraints and takes on a new approach to transforming the original joint\naction-value function into an easily factorizable one, with the same optimal\nactions. QTRAN guarantees more general factorization than VDN or QMIX, thus\ncovering a much wider class of MARL tasks than does previous methods. Our\nexperiments for the tasks of multi-domain Gaussian-squeeze and modified\npredator-prey demonstrate QTRAN's superior performance with especially larger\nmargins in games whose payoffs penalize non-cooperative behavior more\naggressively.\n

Cited by

Related