2022/04/22 by Donghwan Lee, Lee, Donghwan, Do Wan Kim +1
Decision Sciences · Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Innovation Diffusion and Forecasting #Machine Learning (cs.LG) #Systems and Control (eess.SY) #Traffic control and management #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2204.10479
openalex publication_date 2022/04/22 · openalex created_date 2022/04/27 · openalex updated_date 2026/07/28
Temporal difference (TD) learning is a cornerstone reinforcement learning (RL) method for policy evaluation, where the goal is to estimate the value function of a Markov decision process under a fixed policy. While a substantial body of work has established its convergence and stability properties, more recent efforts have focused on its statistical efficiency through finite-time error bounds. In this paper, we advance this line of research by developing a new finite-time error analysis for tabular TD learning that directly exploits a discrete-time stochastic linear system representation and leverages Schur stability of the associated matrices. Beyond the specific bounds obtained, the proposed framework provides a reusable template for analyzing TD learning and related RL algorithms, and it offers control-theoretic insights that may guide future developments in finite-sample RL theory.