English

Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization

Machine Learning 2026-03-10 v1

Abstract

We study infinite-horizon Constrained Markov Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses for constrained reinforcement learning largely rely on tabular policies or linear critics, which limits their applicability to high-dimensional and continuous control problems. We propose a primal-dual natural actor-critic algorithm that integrates neural critic estimation with natural policy gradient updates and leverages Neural Tangent Kernel (NTK) theory to control function-approximation error under Markovian sampling, without requiring access to mixing-time oracles. We establish global convergence and cumulative constraint violation rates of O~(T1/4)\tilde{\mathcal{O}}(T^-1/4) up to approximation errors induced by the policy and critic classes. Our results provide the first such guarantees for CMDPs with general policies and multi-layer neural critics, substantially extending the theoretical foundations of actor-critic methods beyond the linear-critic regime.

Keywords

Cite

@article{arxiv.2603.07698,
  title  = {Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization},
  author = {Anirudh Satheesh and Pankaj Kumar Barman and Washim Uddin Mondal and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2603.07698},
  year   = {2026}
}

Comments

Submitted to UAI 2026