English

On the Convergence of Reinforcement Learning in Nonlinear Continuous State Space Problems

Machine Learning 2021-07-30 v2 Systems and Control Systems and Control

Abstract

We consider the problem of Reinforcement Learning for nonlinear stochastic dynamical systems. We show that in the RL setting, there is an inherent ``Curse of Variance" in addition to Bellman's infamous ``Curse of Dimensionality", in particular, we show that the variance in the solution grows factorial-exponentially in the order of the approximation. A fundamental consequence is that this precludes the search for anything other than ``local" feedback solutions in RL, in order to control the explosive variance growth, and thus, ensure accuracy. We further show that the deterministic optimal control has a perturbation structure, in that the higher order terms do not affect the calculation of lower order terms, which can be utilized in RL to get accurate local solutions.

Keywords

Cite

@article{arxiv.2011.10829,
  title  = {On the Convergence of Reinforcement Learning in Nonlinear Continuous State Space Problems},
  author = {Raman Goyal and Suman Chakravorty and Ran Wang and Mohamed Naveed Gul Mohamed},
  journal= {arXiv preprint arXiv:2011.10829},
  year   = {2021}
}
R2 v1 2026-06-23T20:24:53.255Z