English

On Bellman's Optimality Principle for zs-POSGs

Artificial Intelligence 2022-11-16 v2 Computer Science and Game Theory

Abstract

Many non-trivial sequential decision-making problems are efficiently solved by relying on Bellman's optimality principle, i.e., exploiting the fact that sub-problems are nested recursively within the original problem. Here we show how it can apply to (infinite horizon) 2-player zero-sum partially observable stochastic games (zs-POSGs) by (i) taking a central planner's viewpoint, which can only reason on a sufficient statistic called occupancy state, and (ii) turning such problems into zero-sum occupancy Markov games (zs-OMGs). Then, exploiting the Lipschitz-continuity of the value function in occupancy space, one can derive a version of the HSVI algorithm (Heuristic Search Value Iteration) that provably finds an ϵ\epsilon-Nash equilibrium in finite time.

Keywords

Cite

@article{arxiv.2006.16395,
  title  = {On Bellman's Optimality Principle for zs-POSGs},
  author = {Olivier Buffet and Jilles Dibangoye and Aurélien Delage and Abdallah Saffidine and Vincent Thomas},
  journal= {arXiv preprint arXiv:2006.16395},
  year   = {2022}
}

Comments

18 pages, 0 figures, 1 algorithm

R2 v1 2026-06-23T16:43:03.466Z