2025/02/13 by Max Rudolph, Max Gustav Rudolph, Nathan Lichtlé +18 · 2 voices · 4 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Auction Theory and Applications #Computer science #Economics #Imperfect #Mathematical economics #Perfect information #Philosophy #Simulation Techniques and Applications #cs.LG
paper · pdf · doi:10.48550/arxiv.2502.08938
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/02/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for five large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algorithms for imperfect-information games. Over 7000 training runs, we find that FP-, DO-, and CFR-based approaches fail to outperform generic policy gradient methods. Code is available at https://github.com/nathanlct/IIG-RL-Benchmark and https://github.com/gabrfarina/exp-a-spiel .