vix.ing · top · new · best · stats · spec

Logarithmic Regret in Adaptive Control of Noisy Linear Quadratic Regulator Systems Using Hints

2022/10/28 by Mohammad Esmaeil Akbari, Akbari, Mohammad, Bahman Gharesifard +3
Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #FOS: Mathematics #Optimization and Control (math.OC) #Smart Grid Energy Management

paper · pdf · doi:10.48550/arxiv.2210.16303

openalex publication_date 2022/10/28 · openalex created_date 2022/11/05 · openalex updated_date 2026/07/28

Abstract

The problem of regret minimization for online adaptive control of linear-quadratic systems is studied. In this problem, the true system transition parameters (matrices A and B) are unknown, and the objective is to design and analyze algorithms that generate control policies with sublinear regret. Recent studies show that when the system parameters are fully unknown, there exists a choice of these parameters such that any algorithm that only uses data from the past system trajectory at best achieves a square root of time horizon regret bound, providing a hard fundamental limit on the achievable regret in general. However, it is also known that (poly)-logarithmic regret is achievable when only matrix A or only matrix B is unknown. We present a result, encompassing both scenarios, showing that (poly)-logarithmic regret is achievable when both of these matrices are unknown, but a hint is periodically given to the controller.

Related