2022/03/11 by Min Yin, Yin, Ming, Yaqi Duan +6 · 2 citations
Computer Science · Decision Sciences · Energy · #Advanced Bandit Algorithms Research #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Energy Efficiency and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2203.05804
openalex publication_date 2022/03/11 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
Offline reinforcement learning, which seeks to utilize offline/historical\ndata to optimize sequential decision-making strategies, has gained surging\nprominence in recent studies. Due to the advantage that appropriate function\napproximators can help mitigate the sample complexity burden in modern\nreinforcement learning problems, existing endeavors usually enforce powerful\nfunction representation models (e.g. neural networks) to learn the optimal\npolicies. However, a precise understanding of the statistical limits with\nfunction representations, remains elusive, even when such a representation is\nlinear.\n Towards this goal, we study the statistical limits of offline reinforcement\nlearning with linear model representations. To derive the tight offline\nlearning bound, we design the variance-aware pessimistic value iteration\n(VAPVI), which adopts the conditional variance information of the value\nfunction for time-inhomogeneous episodic linear Markov decision processes\n(MDPs). VAPVI leverages estimated variances of the value functions to reweight\nthe Bellman residuals in the least-square pessimistic value iteration and\nprovides improved offline learning bounds over the best-known existing results\n(whereas the Bellman residuals are equally weighted by design). More\nimportantly, our learning bounds are expressed in terms of system quantities,\nwhich provide natural instance-dependent characterizations that previous\nresults are short of. We hope our results draw a clearer picture of what\noffline learning should look like when linear representations are provided.\n