vix.ing · top · new · best · stats · spec

A Statistical Analysis of Polyak-Ruppert Averaged Q-learning

2021/12/29 by Xiang Li, Wenhao Yang, Li, Xiang +7 · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · #Control Systems and Identification #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2112.14582

openalex publication_date 2021/12/29 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We study Q-learning with Polyak-Ruppert averaging in a discounted Markov decision process in synchronous and tabular settings. Under a Lipschitz condition, we establish a functional central limit theorem for the averaged iteration \boldsymbolQT and show that its standardized partial-sum process converges weakly to a rescaled Brownian motion. The functional central limit theorem implies a fully online inference method for reinforcement learning. Furthermore, we show that \boldsymbolQT is the regular asymptotically linear (RAL) estimator for the optimal Q-value function \boldsymbolQ^* that has the most efficient influence function. We present a nonasymptotic analysis for the ℓ error, 𝔼‖\boldsymbolQT-\boldsymbolQ^*‖, showing that it matches the instance-dependent lower bound for polynomial step sizes. Similar results are provided for entropy-regularized Q-learning without the Lipschitz condition.

Cited by

Related