Markov decision processes: on the convergence of the Monte-Carlo first visit algorithm
Probability
2025-09-23 v2 Optimization and Control
Abstract
We consider the Monte-Carlo first visit algorithm, of which the goal is to find the optimal control in a Markov decision process with finite state space and finite number of possible actions. We show its convergence when the discount factor is smaller than .
Cite
@article{arxiv.2501.08800,
title = {Markov decision processes: on the convergence of the Monte-Carlo first visit algorithm},
author = {Sylvain Delattre and Nicolas Fournier},
journal= {arXiv preprint arXiv:2501.08800},
year = {2025}
}