English

How are policy gradient methods affected by the limits of control?

Optimization and Control 2022-06-15 v1 Machine Learning

Abstract

We study stochastic policy gradient methods from the perspective of control-theoretic limitations. Our main result is that ill-conditioned linear systems in the sense of Doyle inevitably lead to noisy gradient estimates. We also give an example of a class of stable systems in which policy gradient methods suffer from the curse of dimensionality. Our results apply to both state feedback and partially observed systems.

Keywords

Cite

@article{arxiv.2206.06863,
  title  = {How are policy gradient methods affected by the limits of control?},
  author = {Ingvar Ziemann and Anastasios Tsiamis and Henrik Sandberg and Nikolai Matni},
  journal= {arXiv preprint arXiv:2206.06863},
  year   = {2022}
}
R2 v1 2026-06-24T11:50:48.047Z