vix.ing · top · new · best · stats · spec

Elo Ratings for Large Tournaments of Software Agents in Asymmetric Games

2021/04/23 by Ben P. Wise, Wise, Ben
Computer Science · Economics, Econometrics and Finance · Psychology · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Computer Science and Game Theory (cs.GT) #Educational Games and Gamification #FOS: Computer and information sciences #Sports Analytics and Performance

paper · pdf · doi:10.48550/arxiv.2105.00839

openalex publication_date 2021/04/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The Elo rating system has been used world wide for individual sports and team sports, as exemplified by the European Go Federation (EGF), International Chess Federation (FIDE), International Federation of Association Football (FIFA), and many others. To evaluate the performance of artificial intelligence agents, it is natural to evaluate them on the same Elo scale as humans, such as the rating of 5185 attributed to AlphaGo Zero. There are several fundamental differences between humans and AI that suggest modifications to the system, which in turn require revisiting Elo's fundamental rationale. AI is typically trained on many more games than humans play, and we have little a-priori information on newly created AI agents. Further, AI is being extended into games which are asymmetric between the players, and which could even have large complex boards with different setup in every game, such as commercial paper strategy games. We present a revised rating system, and guidelines for tournaments, to reflect these differences.

Related