English

Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization

Machine Learning 2026-02-10 v1

Abstract

We study infinite-horizon average-reward constrained Markov decision processes (CMDPs) under the unichain assumption and general policy parameterizations. Existing regret analyses for constrained reinforcement learning largely rely on ergodicity or strong mixing-time assumptions, which fail to hold in the presence of transient states. We propose a primal--dual natural actor--critic algorithm that leverages multi-level Monte Carlo (MLMC) estimators and an explicit burn-in mechanism to handle unichain dynamics without requiring mixing-time oracles. Our analysis establishes finite-time regret and cumulative constraint violation bounds that scale as O~(T)\tilde{O}(\sqrt{T}), up to approximation errors arising from policy and critic parameterization, thereby extending order-optimal guarantees to a significantly broader class of CMDPs.

Keywords

Cite

@article{arxiv.2602.08000,
  title  = {Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization},
  author = {Anirudh Satheesh and Vaneet Aggarwal},
  journal= {arXiv preprint arXiv:2602.08000},
  year   = {2026}
}
R2 v1 2026-07-01T10:26:48.651Z