English

Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints

Machine Learning 2024-09-27 v2

Abstract

In constrained Markov decision processes (CMDPs) with adversarial rewards and constraints, a well-known impossibility result prevents any algorithm from attaining both sublinear regret and sublinear constraint violation, when competing against a best-in-hindsight policy that satisfies constraints on average. In this paper, we show that this negative result can be eased in CMDPs with non-stationary rewards and constraints, by providing algorithms whose performances smoothly degrade as non-stationarity increases. Specifically, we propose algorithms attaining O~(T+C)\tilde{\mathcal{O}} (\sqrt{T} + C) regret and positive constraint violation under bandit feedback, where CC is a corruption value measuring the environment non-stationarity. This can be Θ(T)\Theta(T) in the worst case, coherently with the impossibility result for adversarial CMDPs. First, we design an algorithm with the desired guarantees when CC is known. Then, in the case CC is unknown, we show how to obtain the same results by embedding such an algorithm in a general meta-procedure. This is of independent interest, as it can be applied to any non-stationary constrained online learning setting.

Keywords

Cite

@article{arxiv.2405.14372,
  title  = {Learning Constrained Markov Decision Processes With Non-stationary Rewards and Constraints},
  author = {Francesco Emanuele Stradi and Anna Lunghi and Matteo Castiglioni and Alberto Marchesi and Nicola Gatti},
  journal= {arXiv preprint arXiv:2405.14372},
  year   = {2024}
}
R2 v1 2026-06-28T16:36:56.531Z