Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation
Machine Learning
2026-05-04 v1 Artificial Intelligence
Optimization and Control
Machine Learning
Abstract
For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We also define the class of multipattern risk-averse problems that generalizes the class of linear systems. We use both concepts in a feature-based -learning method with multipattern -factor approximation and we prove a high-probability regret bound of , where is the horizon, is the mini-batch size, and is the number of episodes. We also propose an economical version of the -learning method that streamlines the policy evaluation (backward) step. The theoretical results are illustrated on a stochastic assignment problem and a short-horizon multi-armed bandit problem.
Cite
@article{arxiv.2605.00654,
title = {Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation},
author = {Andrzej Ruszczynski and Tiangang Zhang},
journal= {arXiv preprint arXiv:2605.00654},
year = {2026}
}