2021/04/15 by Jihao Long, Long, Jihao, Jiequn Han +3
Computer Science · Engineering · #Advancements in Semiconductor Devices and Circuit Design #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2104.07794
openalex publication_date 2021/04/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Reinforcement learning (RL) algorithms based on high-dimensional function approximation have achieved tremendous empirical success in large-scale problems with an enormous number of states. However, most analysis of such algorithms gives rise to error bounds that involve either the number of states or the number of features. This paper considers the situation where the function approximation is made either using the kernel method or the two-layer neural network model, in the context of a fitted Q-iteration algorithm with explicit regularization. We establish an O(H3|\mathcal A|\frac14n-\frac14) bound for the optimal policy with Hn samples, where H is the length of each episode and |\mathcal A| is the size of action space. Our analysis hinges on analyzing the L2 error of the approximated Q-function using n data points. Even though this result still requires a finite-sized action space, the error bound is independent of the dimensionality of the state space.