vix.ing · top · new · best · stats · spec

Finite-Sample Analysis of Off-Policy Natural Actor-Critic with Linear Function Approximation

2021/05/26 by Zaiwei Chen, Chen, Zaiwei, Sajad Khodadadian +3 · 2 citations
Computer Science · Engineering · #Advancements in Semiconductor Devices and Circuit Design #FOS: Computer and information sciences #FOS: Mathematics #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2105.12540

openalex publication_date 2021/05/26 · openalex created_date 2022/10/07 · openalex updated_date 2026/07/28

Abstract

In this paper, we develop a novel variant of off-policy natural actor-critic algorithm with linear function approximation and we establish a sample complexity of O(ε-3), outperforming all the previously known convergence bounds of such algorithms. In order to overcome the divergence due to deadly triad in off-policy policy evaluation under function approximation, we develop a critic that employs n-step TD-learning algorithm with a properly chosen n. We present finite-sample convergence bounds on this critic under both constant and diminishing step sizes, which are of independent interest. Furthermore, we develop a variant of natural policy gradient under function approximation, with an improved convergence rate of O(1/T) after T iterations. Combining the finite sample error bounds of actor and the critic, we obtain the O(ε-3) sample complexity. We derive our sample complexity bounds solely based on the assumption that the behavior policy sufficiently explores all the states and actions, which is a much lighter assumption compared to the related literature.

Cited by

Related