vix.ing · top · new · best · stats · spec

Mathematical methods of reinforcement learning

2026/01/01 by Denis Vital'evich Belomestny, Denis Belomestny, Alexander Vladimirovich Gasnikov +10
Computer Science · #Adaptive Dynamic Programming Control #Bellman equation #Convex analysis #Convex function #Markov chain #Markov decision process #Markov process #Operator (biology) #Q-learning #Reinforcement Learning in Robotics #Reinforcement learning #Stochastic Gradient Optimization Techniques #Variational inequality

paper · doi:10.4213/rm10328

crossref issued 2026/01/01 · crossref published 2026/01/01 · crossref published-print 2026/01/01 · openalex publication_date 2026/01/01 · crossref published-online 2026/07/31 · crossref created 2026/07/31 · crossref deposited 2026/07/31 · crossref indexed 2026/07/31 · openalex created_date 2026/08/01 · openalex updated_date 2026/08/02

Abstract

Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates for value and policy iteration, and temporal-difference schemes. We then develop the optimization perspective: stochastic approximation and martingale methods, convex duality and the role of regularization linking mirror/proximal methods. Function approximation is treated through linear and non-linear settings, covering stabilization, error decomposition, and sample-complexity via concentration inequalities for dependent data and mixing processes. We further cover off-policy evaluation/learning, constrained RL, and constrained MDPs. Throughout, we unify algorithmic templates under common operator and variational lenses, highlighting both finite-sample bounds and asymptotic results. Our presentation is intended to provide a unified mathematical entry point for researchers in probability, optimization, and statistics who are interested in RL. Bibliography: 122 titles.

Citations