English

Mathematical methods of reinforcement learning

Optimization and Control 2026-07-08 v1 Machine Learning

Abstract

Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and the Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates for value and policy iteration, and temporal-difference schemes. We then develop the optimization perspective: stochastic approximation and martingale methods, convex duality and the role of regularization linking mirror/proximal methods. Function approximation is treated through linear and non-linear settings, covering stabilization, error decomposition, and sample-complexity via concentration inequalities for dependent data and mixing processes. We further cover off-policy evaluation/learning, constrained RL and constrained MDPs (CMDPs). Throughout we unify algorithmic templates under common operator and variational lenses, highlighting both finite-sample bounds and asymptotic results. Our presentation is intended to provide a unified mathematical entry point for researchers in probability, optimization, and statistics interested in reinforcement learning.

Cite

@article{arxiv.2607.06935,
  title  = {Mathematical methods of reinforcement learning},
  author = {Denis Belomestny and Alexander Gasnikov and Egor Gladin and Alexey Naumov and Artemy Rubtsov and Yuri Sapronov and Daniil Tiapkin and Nikita Yudin},
  journal= {arXiv preprint arXiv:2607.06935},
  year   = {2026}
}

Comments

65 pages