2019/02/17 by Alon Cohen, Cohen, Alon, Tomer Koren +3 · 6 citations
Computer Science · Decision Sciences · #Adaptive Dynamic Programming Control #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1902.06223
openalex publication_date 2019/02/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present the first computationally-efficient algorithm with widetilde\nO(\√(T)) regret for learning in Linear Quadratic Control systems with\nunknown dynamics. By that, we resolve an open question of Abbasi-Yadkori and\nSzepesv 'ari (2011) and Dean, Mania, Matni, Recht, and Tu (2018).\n