2021/02/15 by Wei-Fang Sun, Cheng‐Kuang Lee, Sun, Wei-Fang +5 · 2 citations
Computer Science · Mathematics · #Adaptive Dynamic Programming Control #Algorithm #Applied mathematics #Artificial intelligence #Bellman equation #Computer science #Econometrics #Evolutionary Algorithms and Applications #Expected value #FOS: Computer and information sciences #Factorization #Function (biology) #Machine Learning (cs.LG) #Machine learning #Mathematical optimization #Mathematics #Moment-generating function #Multiagent Systems (cs.MA) #Observability #Quantile #Quantile function #Random variable #Reinforcement Learning in Robotics #Reinforcement learning #Statistics #Value (mathematics) #cs.LG #cs.MA
paper · pdf · doi:10.48550/arxiv.2102.07936
published in arXiv (Cornell University) (Cornell University) · ICML 2021
openalex publication_date 2021/02/15 · openalex created_date 2021/06/22 · arxiv created 2021/12/22 · arxiv updated 2021/12/23 · openalex updated_date 2026/08/05
In fully cooperative multi-agent reinforcement learning (MARL) settings, the\nenvironments are highly stochastic due to the partial observability of each\nagent and the continuously changing policies of the other agents. To address\nthe above issues, we integrate distributional RL and value function\nfactorization methods by proposing a Distributional Value Function\nFactorization (DFAC) framework to generalize expected value function\nfactorization methods to their DFAC variants. DFAC extends the individual\nutility functions from deterministic variables to random variables, and models\nthe quantile function of the total return as a quantile mixture. To validate\nDFAC, we demonstrate DFAC's ability to factorize a simple two-step matrix game\nwith stochastic rewards and perform experiments on all Super Hard tasks of\nStarCraft Multi-Agent Challenge, showing that DFAC is able to outperform\nexpected value function factorization baselines.\n