English

Global Convergence of Policy Gradient Methods for ReLU Controllers in Linear Quadratic Regulation

Optimization and Control 2026-04-27 v1 Systems and Control Systems and Control

Abstract

We study the convergence of model-based policy gradient for the deterministic, scalar, discounted linear-quadratic regulator when the controller is an overparameterized one-hidden-layer ReLU network without biases. Although the optimal LQR controller is linear, neural parameterization creates a redundant nonconvex weight space with a possibly asymmetric piecewise-linear controller. We show that this structure can still be analyzed exactly through the two effective gains induced on the positive and negative half-lines. Under suitable random initialization, sufficient width, and a small step size, the model-based policy gradient remains stable, decreases the cost geometrically, and drives the effective gains to the unique optimal scalar LQR gain with high probability.

Keywords

Cite

@article{arxiv.2604.22138,
  title  = {Global Convergence of Policy Gradient Methods for ReLU Controllers in Linear Quadratic Regulation},
  author = {Jhojan A. Rodriguez-Gil and César A. Uribe},
  journal= {arXiv preprint arXiv:2604.22138},
  year   = {2026}
}