English

Convergence of Fast Policy Iteration in Markov Games and Robust MDPs

Computer Science and Game Theory 2025-11-18 v2

Abstract

Markov games and robust MDPs are closely related models that involve computing a pair of saddle point policies. As part of the long-standing effort to develop efficient algorithms for these models, the Filar-Tolwinski (FT) algorithm has shown considerable promise. As our first contribution, we demonstrate that FT may fail to converge to a saddle point and may loop indefinitely, even in small games. This observation contradicts the proof of FT's convergence to a saddle point in the original paper. As our second contribution, we propose Residual Conditioned Policy Iteration (RCPI). RCPI builds on FT, but is guaranteed to converge to a saddle point. Our numerical results show that RCPI outperforms other convergent algorithms by several orders of magnitude.

Keywords

Cite

@article{arxiv.2508.06661,
  title  = {Convergence of Fast Policy Iteration in Markov Games and Robust MDPs},
  author = {Keith Badger and Jefferson Huang and Marek Petrik},
  journal= {arXiv preprint arXiv:2508.06661},
  year   = {2025}
}