vix.ing · top · new · best · stats · spec

Policy Optimization for \H2 Linear Control with\n \H_\∞ Robustness Guarantee: Implicit Regularization and Global\n Convergence

2019/10/21 by Kaiqing Zhang, Bin Hu, Zhang, Kaiqing +3 · 3 citations
Computer Science · Engineering · #Adaptive Dynamic Programming Control #FOS: Computer and information sciences #FOS: Electrical engineering #FOS: Mathematics #Machine Learning (cs.LG) #Mechanical Circulatory Support Devices #Optimization and Control (math.OC) #Systems and Control (eess.SY) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1910.09496

openalex publication_date 2019/10/21 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28

Abstract

Policy optimization (PO) is a key ingredient for reinforcement learning (RL).\nFor control design, certain constraints are usually enforced on the policies to\noptimize, accounting for either the stability, robustness, or safety concerns\non the system. Hence, PO is by nature a constrained (nonconvex) optimization in\nmost cases, whose global convergence is challenging to analyze in general. More\nimportantly, some constraints that are safety-critical, e.g., the\n\H_\∞-norm constraint that guarantees the system robustness, are\ndifficult to enforce as the PO methods proceed. Recently, policy gradient\nmethods have been shown to converge to the global optimum of linear quadratic\nregulator (LQR), a classical optimal control problem, without\nregularizing/projecting the control iterates onto the stabilizing set, its\n(implicit) feasible set. This striking result is built upon the coercive\nproperty of the cost, ensuring that the iterates remain feasible as the cost\ndecreases. In this paper, we study the convergence theory of PO for\n\H2 linear control with \H_\∞-norm robustness\nguarantee. One significant new feature of this problem is the lack of\ncoercivity, i.e., the cost may have finite value around the feasible set\nboundary, breaking the existing analysis for LQR. Interestingly, we show that\ntwo PO methods enjoy the implicit regularization property, i.e., the iterates\npreserve the \H_\∞ robustness constraint as if they are\nregularized by the algorithms. Furthermore, despite the nonconvexity of the\nproblem, we show that these algorithms converge to the globally optimal\npolicies with globally sublinear rates, avoiding all suboptimal stationary\npoints/local minima, and with locally (super-)linear rates under certain\nconditions.\n

Cited by

Related