2020/10/21 by Deepan Muthirayan, Pramod P. Khargonekar, Muthirayan, Deepan +1
Computer Science · Decision Sciences · Engineering · #Adaptive Dynamic Programming Control #Advanced Bandit Algorithms Research #Advanced Control Systems Optimization #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2010.11327
openalex publication_date 2020/10/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper we provide provable regret guarantees for an online meta-learning receding horizon control algorithm in an iterative control setting. We consider the setting where, in each iteration the system to be controlled is a linear deterministic system that is different and unknown, the cost for the controller in an iteration is a general additive cost function and there are affine control input constraints. By analysing conditions under which sub-linear regret is achievable, we prove that the meta-learning online receding horizon controller achieves an average of the dynamic regret for the controller cost that is O((1+1/√(N))T3/4) with the number of iterations N. Thus, we show that the worst regret for learning within an iteration improves with experience of more iterations, with guarantee on rate of improvement.