English

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

Machine Learning 2026-07-25 v1 Optimization and Control Machine Learning

Abstract

Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success. In this paper, we study exact NPG in finite-horizon Markov Decision Processes with known dynamics and horizon-dependent transition kernels. We provide the first finite-time convergence guarantees for this algorithm in this setting, for which we consider both constant and increasing step size regimes. With a constant step size ηt=η\eta_t=\eta, we prove that NPG converges sublinearly with a rate of O(H2/t)\mathcal{O}(H^{2}/t) after tt iterations, where HH is the horizon length. We also extend this constant step size analysis to linear MDPs in an exact population-projection oracle under a full support projection distribution, recovering the same sublinear rate as in the tabular setting. Furthermore, with increasing step sizes, we prove that this algorithm achieves a linear convergence rate of O((11ϑρ)t)\mathcal{O}\left(\left(1-\frac{1}{\vartheta_\rho}\right)^t\right) for a problem-dependent constant ϑρ>1\vartheta_\rho > 1, and the horizon-only robust schedule of the form ηt=η0(H/(H1))t\eta_t=\eta_0(H/(H-1))^t where η0>0\eta_0>0 and H2H \geq 2, attains this same geometric rate.

Keywords

Cite

@article{arxiv.2607.22982,
  title  = {Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes},
  author = {Asha Barua and Sajad Khodadadian},
  journal= {arXiv preprint arXiv:2607.22982},
  year   = {2026}
}

Comments

33 pages, 3 figures (each with two subfigures), 1 table. Accepted at the Reinforcement Learning Conference (RLC 2026)