English

On Linear Convergence of Policy Gradient Methods for Finite MDPs

Machine Learning 2021-12-14 v2 Optimization and Control Machine Learning

Abstract

We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic policies and with exact gradient evaluations. There has been some recent work viewing this setting as an instance of smooth non-linear optimization problems and showing sub-linear convergence rates with small step-sizes. Here, we take a different perspective based on connections with policy iteration and show that many variants of policy gradient methods succeed with large step-sizes and attain a linear rate of convergence.

Keywords

Cite

@article{arxiv.2007.11120,
  title  = {On Linear Convergence of Policy Gradient Methods for Finite MDPs},
  author = {Jalaj Bhandari and Daniel Russo},
  journal= {arXiv preprint arXiv:2007.11120},
  year   = {2021}
}

Comments

Published in AISTATS 2021