vix.ing · top · new · best · stats · spec

Limiting dynamics for Q-learning with memory one in symmetric two-player, two-action games

2021/07/29 by Janusz M Meylahn, Meylahn, Janusz M, Janssen, Lars
Biochemistry, Genetics and Molecular Biology · Decision Sciences · #Adaptation and Self-Organizing Systems (nlin.AO) #Dynamical Systems (math.DS) #FOS: Mathematics #FOS: Physical sciences #Game Theory and Applications #Receptor Mechanisms and Signaling

paper · pdf · doi:10.48550/arxiv.2107.13995

openalex publication_date 2021/07/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We develop a method based on computer algebra systems to represent the mutual pure strategy best-response dynamics of symmetric two-player, two-action repeated games played by players with a one-period memory. We apply this method to the iterated prisoner's dilemma, stag hunt and hawk-dove games and identify all possible equilibrium strategy pairs and the conditions for their existence. The only equilibrium strategy pair that is possible in all three games is the win-stay, lose-shift strategy. Lastly, we show that the mutual best-response dynamics are realized by a sample batch Q-learning algorithm in the infinite batch size limit.

Related