2020/03/12 by Sahin Lale, Lale, Sahin, Kamyar Azizzadenesheli +5
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Advanced Control Systems Optimization #FOS: Computer and information sciences #FOS: Mathematics #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Control (math.OC)
paper · pdf · doi:10.48550/arxiv.2003.05999
openalex publication_date 2020/03/12 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
We study the problem of adaptive control in partially observable linear\nquadratic Gaussian control systems, where the model dynamics are unknown a\npriori. We propose LqgOpt, a novel reinforcement learning algorithm based on\nthe principle of optimism in the face of uncertainty, to effectively minimize\nthe overall control cost. We employ the predictor state evolution\nrepresentation of the system dynamics and deploy a recently proposed\nclosed-loop system identification method, estimation, and confidence bound\nconstruction. LqgOpt efficiently explores the system dynamics, estimates the\nmodel parameters up to their confidence interval, and deploys the controller of\nthe most optimistic model for further exploration and exploitation. We provide\nstability guarantees for LqgOpt and prove the regret upper bound of\n\\O(\√(T)) for adaptive control of linear quadratic\nGaussian (LQG) systems, where T is the time horizon of the problem.\n