English

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

Machine Learning 2026-02-04 v1 Artificial Intelligence

Abstract

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depends on entire trajectories without admitting a recursive Bellman decomposition in the underlying Markov decision process. A classical resolution relies on state augmentation with a continuous variable. However, unless restricted to a specialized class of admissible value functions, this formulation induces sparse rewards and degenerate fixed points. In this work, we propose a novel formulation of the static CVaR objective based on augmentation. Our alternative approach leads to a Bellman operator with: (1) dense per-step rewards; (2) contracting properties on the full space of bounded value functions. Building on this theoretical foundation, we develop risk-averse value iteration and model-free Q-learning algorithms that rely on discretized augmented states. We further provide convergence guarantees and approximation error bounds due to discretization. Empirical results demonstrate that our algorithms successfully learn CVaR-sensitive policies and achieve effective performance-safety trade-offs.

Cite

@article{arxiv.2602.03778,
  title  = {Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity},
  author = {Aneri Muni and Vincent Taboga and Esther Derman and Pierre-Luc Bacon and Erick Delage},
  journal= {arXiv preprint arXiv:2602.03778},
  year   = {2026}
}
R2 v1 2026-07-01T09:34:42.569Z