vix.ing · top · new · best · stats · spec

A Q-learning algorithm for discrete-time linear-quadratic control with random parameters of unknown distribution: convergence and stabilization

2020/11/10 by Kai Du, Qingxin Meng, Du, Kai +3 · 1 citation
Computer Science · Engineering · #49N10 #93D15 #93E35 #Adaptive Dynamic Programming Control #Advanced Control Systems Optimization #Control Systems and Identification #FOS: Mathematics #Optimization and Control (math.OC) #Probability (math.PR)

paper · pdf · doi:10.48550/arxiv.2011.04970

openalex publication_date 2020/11/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper studies an infinite horizon optimal control problem for discrete-time linear systems and quadratic criteria, both with random parameters which are independent and identically distributed with respect to time. A classical approach is to solve an algebraic Riccati equation that involves mathematical expectations and requires certain statistical information of the parameters. In this paper, we propose an online iterative algorithm in the spirit of Q-learning for the situation where only one random sample of parameters emerges at each time step. The first theorem proves the equivalence of three properties: the convergence of the learning sequence, the well-posedness of the control problem, and the solvability of the algebraic Riccati equation. The second theorem shows that the adaptive feedback control in terms of the learning sequence stabilizes the system as long as the control problem is well-posed. Numerical examples are presented to illustrate our results.

Cited by

Related