2024/05/28 by Johann Bauer, S West, Bauer, Johann +5
Computer Science · Economics, Econometrics and Finance · Social Sciences · #37N40 (Primary) 91A26 (Secondary) #Artificial Intelligence in Games #Dynamical Systems (math.DS) #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Optimization and Control (math.OC) #Populations and Evolution (q-bio.PE) #Sports Analytics and Performance #Wikis in Education and Collaboration
paper · doi:10.48550/arxiv.2405.18190
openalex publication_date 2024/05/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present two variants of a multi-agent reinforcement learning algorithm based on evolutionary game theoretic considerations. The intentional simplicity of one variant enables us to prove results on its relationship to a system of ordinary differential equations of replicator-mutator dynamics type, allowing us to present proofs on the algorithm's convergence conditions in various settings via its ODE counterpart. The more complicated variant enables comparisons to Q-learning based algorithms. We compare both variants experimentally to WoLF-PHC and frequency-adjusted Q-learning on a range of settings, illustrating cases of increasing dimensionality where our variants preserve convergence in contrast to more complicated algorithms. The availability of analytic results provides a degree of transferability of results as compared to purely empirical case studies, illustrating the general utility of a dynamical systems perspective on multi-agent reinforcement learning when addressing questions of convergence and reliable generalisation.