English

Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall

Machine Learning 2021-06-14 v1 Machine Learning

Abstract

We study the problem of learning a Nash equilibrium (NE) in an imperfect information game (IIG) through self-play. Precisely, we focus on two-player, zero-sum, episodic, tabular IIG under the perfect-recall assumption where the only feedback is realizations of the game (bandit feedback). In particular, the dynamic of the IIG is not known -- we can only access it by sampling or interacting with a game simulator. For this learning setting, we provide the Implicit Exploration Online Mirror Descent (IXOMD) algorithm. It is a model-free algorithm with a high-probability bound on the convergence rate to the NE of order 1/T1/\sqrt{T} where TT is the number of played games. Moreover, IXOMD is computationally efficient as it needs to perform the updates only along the sampled trajectory.

Cite

@article{arxiv.2106.06279,
  title  = {Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall},
  author = {Tadashi Kozuno and Pierre Ménard and Rémi Munos and Michal Valko},
  journal= {arXiv preprint arXiv:2106.06279},
  year   = {2021}
}

Comments

20 pages

R2 v1 2026-06-24T03:05:39.045Z