English

Non-myopic learning in repeated stochastic games

Computer Science and Game Theory 2018-01-22 v3 Artificial Intelligence Machine Learning

Abstract

In repeated stochastic games (RSGs), an agent must quickly adapt to the behavior of previously unknown associates, who may themselves be learning. This machine-learning problem is particularly challenging due, in part, to the presence of multiple (even infinite) equilibria and inherently large strategy spaces. In this paper, we introduce a method to reduce the strategy space of two-player general-sum RSGs to a handful of expert strategies. This process, called Mega, effectually reduces an RSG to a bandit problem. We show that the resulting strategy space preserves several important properties of the original RSG, thus enabling a learner to produce robust strategies within a reasonably small number of interactions. To better establish strengths and weaknesses of this approach, we empirically evaluate the resulting learning system against other algorithms in three different RSGs.

Keywords

Cite

@article{arxiv.1409.8498,
  title  = {Non-myopic learning in repeated stochastic games},
  author = {Jacob W. Crandall},
  journal= {arXiv preprint arXiv:1409.8498},
  year   = {2018}
}
R2 v1 2026-06-22T06:09:22.569Z