English

Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs

Machine Learning 2026-05-04 v2 Artificial Intelligence

Abstract

This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regularized Accelerated Natural Policy Gradient (PDR-ANPG) algorithm that uses entropy and quadratic regularizers to reach this goal. For parameterized policy classes with a transferred compatibility approximation error, ϵbias\epsilon_{\mathrm{bias}}, PDR-ANPG achieves a last-iterate ϵ\epsilon optimality gap and ϵ\epsilon constraint violation with a sample complexity of O~(ϵ2min{ϵ2,ϵbias13})\tilde{\mathcal{O}}(\epsilon^{-2}\min\{\epsilon^{-2},\epsilon_{\mathrm{bias}}^{-\frac{1}{3}}\}). If the class is incomplete (ϵbias>0\epsilon_{\mathrm{bias}}>0), then the sample complexity reduces to O~(ϵ2)\tilde{\mathcal{O}}(\epsilon^{-2}) for ϵ<(ϵbias)16\epsilon<(\epsilon_{\mathrm{bias}})^{\frac{1}{6}}. Moreover, for complete policies with ϵbias=0\epsilon_{\mathrm{bias}}=0, our algorithm achieves a last-iterate ϵ\epsilon optimality gap and ϵ\epsilon constraint violation with O~(ϵ4)\tilde{\mathcal{O}}(\epsilon^{-4}) sample complexity. It is a significant improvement over the state-of-the-art last-iterate guarantees of general parameterized CMDPs.

Keywords

Cite

@article{arxiv.2408.11513,
  title  = {Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs},
  author = {Washim Uddin Mondal and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2408.11513},
  year   = {2026}
}

Comments

Published in Transactions on Machine Learning Research (TMLR)