2020/10/21 by Tampubolon, Ezra, Ceribasic, Haris, Boche, Holger
#Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #FOS: Economics and business #FOS: Electrical engineering #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Systems and Control (eess.SY) #Theoretical Economics (econ.TH) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.2010.10901
In this work, we study the system of interacting non-cooperative two Q-learning agents, where one agent has the privilege of observing the other's actions. We show that this information asymmetry can lead to a stable outcome of population learning, which generally does not occur in an environment of general independent learners. The resulting post-learning policies are almost optimal in the underlying game sense, i.e., they form a Nash equilibrium. Furthermore, we propose in this work a Q-learning algorithm, requiring predictive observation of two subsequent opponent's actions, yielding an optimal strategy given that the latter applies a stationary strategy, and discuss the existence of the Nash equilibrium in the underlying information asymmetrical game.