English

On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents

Machine Learning 2026-05-22 v1 Systems and Control Systems and Control Machine Learning

Abstract

We study risk-sensitive reinforcement learning in finite discounted MDPs, where a generative model of the MDP is assumed to be available. We consider a family or risk measures called the optimized certainty equivalent (OCE), which includes important risk measures such as entropic risk, CVaR, and mean-variance. Our focus is on the sample complexities of learning the optimal state-action value function (value learning) and an optimal policy (policy learning) under recursive OCE. We provide an exact characterization of utility functions uu for which the corresponding OCE defines an objective that is PAC-learnable. We analyze a simple model-based approach and derive PAC sample complexity bounds. We establish that whenever uu does not have full domain dom(u)R\text{dom}(u)\neq \mathbb{R}, the corresponding problem is not PAC-learnable. Finally, we establish corresponding lower bounds for both value and policy learning, demonstrating tightness in the size SASA of state-action space, and for a more restricted class of utilities, we derive lower bounds that makes the dependence on the effective horizon 11γ\frac{1}{1-\gamma} explicit. Specifically, for CVaRτ\text{CVaR}_\tau we show that the correct dependence on τ\tau is 1τ2\frac{1}{\tau^2}, thus improving by a factor of 1τ\frac{1}{\tau} over state-of-the-art although our bound has a suboptimal dependence on 11γ\frac{1}{1-\gamma}.

Keywords

Cite

@article{arxiv.2605.21763,
  title  = {On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents},
  author = {Oliver Mortensen and Mohammad Sadegh Talebi},
  journal= {arXiv preprint arXiv:2605.21763},
  year   = {2026}
}

Comments

Accepted to RLC 2026. arXiv admin note: substantial text overlap with arXiv:2506.00286

R2 v1 2026-07-22T07:24:59.456Z