Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases
Abstract
While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empirical behavior. In the discounted-reward case, classical theory guarantees geometric convergence with rate , while in the average-reward case recent work suggests that only sublinear convergence can be expected. In practice, however, VI is often observed to converge significantly faster. In this work, we show through a unified geometry-based analysis that, under an assumption of a unique and unichain optimal policy, (i) convergence is geometric in both the discounted- and average-reward settings and (ii) the convergence rate is faster than previous analyses suggest.
Cite
@article{arxiv.2510.23914,
title = {Revisiting Value Iteration: Unified Analysis of Discounted and Average-Reward Cases},
author = {Arsenii Mustafin and Xinyi Sheng and Dominik Baumann},
journal= {arXiv preprint arXiv:2510.23914},
year = {2026}
}
Comments
22 pages, 2 figure