2020/05/22 by Arta Seify, Michael Buro, Seify, Arta +1
Computer Science · Economics, Econometrics and Finance · Psychology · #Artificial Intelligence in Games #Sports Analytics and Performance #Gambling Behavior and Treatments
paper · pdf · doi:10.48550/arxiv.2005.11335
The combination of Monte-Carlo Tree Search (MCTS) and deep reinforcement\nlearning is state-of-the-art in two-player perfect-information games. In this\npaper, we describe a search algorithm that uses a variant of MCTS which we\nenhanced by 1) a novel action value normalization mechanism for games with\npotentially unbounded rewards (which is the case in many optimization\nproblems), 2) defining a virtual loss function that enables effective search\nparallelization, and 3) a policy network, trained by generations of self-play,\nto guide the search. We gauge the effectiveness of our method in "SameGame"---a\npopular single-player test domain. Our experimental results indicate that our\nmethod outperforms baseline algorithms on several board sizes. Additionally, it\nis competitive with state-of-the-art search algorithms on a public set of\npositions.\n