vix.ing · top · new · best · stats · spec

Escaping High-order Saddles in Policy Optimization for Linear Quadratic Gaussian (LQG) Control

2022/04/02 by Yang Zheng, Yue Sun, Zheng, Yang +5 · 2 citations
Computer Science · #Adaptive Dynamic Programming Control #Age of Information Optimization #Dynamical Systems (math.DS) #FOS: Electrical engineering #FOS: Mathematics #Optimization and Control (math.OC) #Reinforcement Learning in Robotics #Systems and Control (eess.SY) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2204.00912

openalex publication_date 2022/04/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

First order policy optimization has been widely used in reinforcement learning. It guarantees to find the optimal policy for the state-feedback linear quadratic regulator (LQR). However, the performance of policy optimization remains unclear for the linear quadratic Gaussian (LQG) control where the LQG cost has spurious suboptimal stationary points. In this paper, we introduce a novel perturbed policy gradient (PGD) method to escape a large class of bad stationary points (including high-order saddles). In particular, based on the specific structure of LQG, we introduce a novel reparameterization procedure which converts the iterate from a high-order saddle to a strict saddle, from which standard random perturbations in PGD can escape efficiently. We further characterize the high-order saddles that can be escaped by our algorithm.

Cited by

Related