This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regularized Accelerated Natural Policy Gradient (PDR-ANPG) algorithm that uses entropy and quadratic regularizers to reach this goal. For parameterized policy classes with a transferred compatibility approximation error, ϵbias, PDR-ANPG achieves a last-iterate ϵ optimality gap and ϵ constraint violation with a sample complexity of O~(ϵ−2min{ϵ−2,ϵbias−31}). If the class is incomplete (ϵbias>0), then the sample complexity reduces to O~(ϵ−2) for ϵ<(ϵbias)61. Moreover, for complete policies with ϵbias=0, our algorithm achieves a last-iterate ϵ optimality gap and ϵ constraint violation with O~(ϵ−4) sample complexity. It is a significant improvement over the state-of-the-art last-iterate guarantees of general parameterized CMDPs.
@article{arxiv.2408.11513,
title = {Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs},
author = {Washim Uddin Mondal and Vaneet Aggarwal},
journal= {arXiv preprint arXiv:2408.11513},
year = {2026}
}
Comments
Published in Transactions on Machine Learning Research (TMLR)