2022/06/14 by Ingvar Ziemann, Anastasios Tsiamis, Ziemann, Ingvar +5
Mathematics · #Markov Chains and Monte Carlo Methods
paper · pdf · doi:10.48550/arxiv.2206.06863
We study stochastic policy gradient methods from the perspective of control-theoretic limitations. Our main result is that ill-conditioned linear systems in the sense of Doyle inevitably lead to noisy gradient estimates. We also give an example of a class of stable systems in which policy gradient methods suffer from the curse of dimensionality. Our results apply to both state feedback and partially observed systems.