vix.ing · top · new · best · stats

A study of first-passage time minimization via Q-learning in heated gridworlds

2021/10/05 by M. A. Larchenko, Maria Larchenko, Larchenko, M. A. +9
Computer Science · Engineering · Mathematics · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Distributed Control Multi-Agent Systems #Dynamical Systems (math.DS) #FOS: Computer and information sciences #FOS: Mathematics #FOS: Physical sciences #Machine Learning (cs.LG) #Modular Robots and Swarm Intelligence #Optimization and Control (math.OC) #Optimization and Search Problems #Statistical Mechanics (cond-mat.stat-mech) #cond-mat.stat-mech #cs.AI #cs.LG #math.DS #math.OC

paper · pdf · doi:10.48550/arxiv.2110.02129

arxiv created 2021/10/05 · openalex publication_date 2021/10/05 · arxiv updated 2021/10/06 · openalex created_date 2021/10/11 · openalex updated_date 2026/07/28

Abstract

Optimization of first-passage times is required in applications ranging from nanobots navigation to market trading. In such settings, one often encounters unevenly distributed noise levels across the environment. We extensively study how a learning agent fares in 1- and 2- dimensional heated gridworlds with an uneven temperature distribution. The results show certain bias effects in agents trained via simple tabular Q-learning, SARSA, Expected SARSA and Double Q-learning. While high learning rate prevents exploration of regions with higher temperature, low enough rate increases the presence of agents in such regions. The discovered peculiarities and biases of temporal-difference-based reinforcement learning methods should be taken into account in real-world physical applications and agent design.

Citations

Related