2024/02/07 by Ki-Hyuk Hong, Ambuj Tewari, Hong, Kihyuk +1 · 2 citations
Computer Science · Engineering · #Elevator Systems and Control #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Smart Parking Systems Research
paper · pdf · doi:10.48550/arxiv.2402.04493
openalex publication_date 2024/02/07 · openalex created_date 2024/02/09 · openalex updated_date 2026/07/28
We study offline reinforcement learning (RL) with linear MDPs under the infinite-horizon discounted setting which aims to learn a policy that maximizes the expected discounted cumulative reward using a pre-collected dataset. Existing algorithms for this setting either require a uniform data coverage assumptions or are computationally inefficient for finding an ε-optimal policy with O(ε-2) sample complexity. In this paper, we propose a primal dual algorithm for offline RL with linear MDPs in the infinite-horizon discounted setting. Our algorithm is the first computationally efficient algorithm in this setting that achieves sample complexity of O(ε-2) with partial data coverage assumption. Our work is an improvement upon a recent work that requires O(ε-4) samples. Moreover, we extend our algorithm to work in the offline constrained RL setting that enforces constraints on additional reward signals.