English

Analysis of Lower Bounds for Simple Policy Iteration

Machine Learning 2019-12-02 v1 Optimization and Control Machine Learning

Abstract

Policy iteration is a family of algorithms that are used to find an optimal policy for a given Markov Decision Problem (MDP). Simple Policy iteration (SPI) is a type of policy iteration where the strategy is to change the policy at exactly one improvable state at every step. Melekopoglou and Condon [1990] showed an exponential lower bound on the number of iterations taken by SPI for a 2 action MDP. The results have not been generalized to kk-action MDP since. In this paper, we revisit the algorithm and the analysis done by Melekopoglou and Condon. We generalize the previous result and prove a novel exponential lower bound on the number of iterations taken by policy iteration for NN-state, kk-action MDPs. We construct a family of MDPs and give an index-based switching rule that yields a strong lower bound of O((3+k)2N/23)\mathcal{O}\big((3+k)2^{N/2-3}\big).

Keywords

Cite

@article{arxiv.1911.12842,
  title  = {Analysis of Lower Bounds for Simple Policy Iteration},
  author = {Sarthak Consul and Bhishma Dedhia and Kumar Ashutosh and Parthasarathi Khirwadkar},
  journal= {arXiv preprint arXiv:1911.12842},
  year   = {2019}
}
R2 v1 2026-06-23T12:30:25.753Z