English

Uniform Value and Decidability in Ergodic Blind Stochastic Games

Optimization and Control 2025-11-24 v2 Computational Complexity

Abstract

We study a class of two-player zero-sum stochastic games known as \textit{blind stochastic games}, where players neither observe the state nor receive any information about it during the game. A central concept for analyzing long-duration stochastic games is the \textit{uniform value}. A game has a uniform value vv if for every ε>0\varepsilon>0, Player 1 (resp., Player 2) has a strategy such that, for all sufficiently large nn, his average payoff over nn stages is at least vεv-\varepsilon (resp., at most v+εv+\varepsilon). Prior work has shown that the uniform value may not exist in general blind stochastic games. To address this, we introduce a subclass called \textit{ergodic blind stochastic games}, defined by imposing an ergodicity condition on the state transitions. For this subclass, we prove the existence of the uniform value and provide an algorithm to approximate it, establishing the \textit{decidability} of the approximation problem. Notably, this decidability result is novel even in the single-player setting of Partially Observable Markov Decision Processes (POMDPs). Furthermore, we show that no algorithm can compute the uniform value exactly, emphasizing the tightness of our result. Finally, we establish that the uniform value is independent of the initial belief.

Keywords

Cite

@article{arxiv.2405.12583,
  title  = {Uniform Value and Decidability in Ergodic Blind Stochastic Games},
  author = {Krishnendu Chatterjee and David Lurie and Raimundo Saona and Bruno Ziliotto},
  journal= {arXiv preprint arXiv:2405.12583},
  year   = {2025}
}
R2 v1 2026-06-28T16:33:58.959Z