On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
Abstract
We study risk-sensitive reinforcement learning in finite discounted MDPs, where a generative model of the MDP is assumed to be available. We consider a family or risk measures called the optimized certainty equivalent (OCE), which includes important risk measures such as entropic risk, CVaR, and mean-variance. Our focus is on the sample complexities of learning the optimal state-action value function (value learning) and an optimal policy (policy learning) under recursive OCE. We provide an exact characterization of utility functions for which the corresponding OCE defines an objective that is PAC-learnable. We analyze a simple model-based approach and derive PAC sample complexity bounds. We establish that whenever does not have full domain , the corresponding problem is not PAC-learnable. Finally, we establish corresponding lower bounds for both value and policy learning, demonstrating tightness in the size of state-action space, and for a more restricted class of utilities, we derive lower bounds that makes the dependence on the effective horizon explicit. Specifically, for we show that the correct dependence on is , thus improving by a factor of over state-of-the-art although our bound has a suboptimal dependence on .
Cite
@article{arxiv.2605.21763,
title = {On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents},
author = {Oliver Mortensen and Mohammad Sadegh Talebi},
journal= {arXiv preprint arXiv:2605.21763},
year = {2026}
}
Comments
Accepted to RLC 2026. arXiv admin note: substantial text overlap with arXiv:2506.00286