English

Upper Bounds for All and Max-gain Policy Iteration Algorithms on Deterministic MDPs

Discrete Mathematics 2023-10-10 v2 Computational Complexity Combinatorics

Abstract

Policy Iteration (PI) is a widely used family of algorithms to compute optimal policies for Markov Decision Problems (MDPs). We derive upper bounds on the running time of PI on Deterministic MDPs (DMDPs): the class of MDPs in which every state-action pair has a unique next state. Our results include a non-trivial upper bound that applies to the entire family of PI algorithms; another to all "max-gain" switching variants; and affirmation that a conjecture regarding Howard's PI on MDPs is true for DMDPs. Our analysis is based on certain graph-theoretic results, which may be of independent interest.

Keywords

Cite

@article{arxiv.2211.15602,
  title  = {Upper Bounds for All and Max-gain Policy Iteration Algorithms on Deterministic MDPs},
  author = {Ritesh Goenka and Eashan Gupta and Sushil Khyalia and Pratyush Agarwal and Mulinti Shaik Wajid and Shivaram Kalyanakrishnan},
  journal= {arXiv preprint arXiv:2211.15602},
  year   = {2023}
}

Comments

Added new bounds for two state MDPs

R2 v1 2026-06-28T07:15:24.746Z