2020/06/29 by Shuyue Hu, Chin-Wing Leung, Hu, Shuyue +5
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Evolution and Genetic Dynamics #FOS: Computer and information sciences #Game Theory and Applications #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2006.16068
openalex publication_date 2020/06/29 · openalex created_date 2020/07/02 · openalex updated_date 2026/07/28
Understanding the evolutionary dynamics of reinforcement learning under multi-agent settings has long remained an open problem. While previous works primarily focus on 2-player games, we consider population games, which model the strategic interactions of a large population comprising small and anonymous agents. This paper presents a formal relation between stochastic processes and the dynamics of independent learning agents who reason based on the reward signals. Using a master equation approach, we provide a novel unified framework for characterising population dynamics via a single partial differential equation (Theorem 1). Through a case study involving Cross learning agents, we illustrate that Theorem 1 allows us to identify qualitatively different evolutionary dynamics, to analyse steady states, and to gain insights into the expected behaviour of a population. In addition, we present extensive experimental results validating that Theorem 1 holds for a variety of learning methods and population games.