English

Maximin Action Identification: A New Bandit Framework for Games

Statistics Theory 2016-02-16 v1 Computer Science and Game Theory Machine Learning Statistics Theory

Abstract

We study an original problem of pure exploration in a strategic bandit model motivated by Monte Carlo Tree Search. It consists in identifying the best action in a game, when the player may sample random outcomes of sequentially chosen pairs of actions. We propose two strategies for the fixed-confidence setting: Maximin-LUCB, based on lower-and upper-confidence bounds; and Maximin-Racing, which operates by successively eliminating the sub-optimal actions. We discuss the sample complexity of both methods and compare their performance empirically. We sketch a lower bound analysis, and possible connections to an optimal algorithm.

Cite

@article{arxiv.1602.04676,
  title  = {Maximin Action Identification: A New Bandit Framework for Games},
  author = {Aurélien Garivier and Emilie Kaufmann and Wouter Koolen},
  journal= {arXiv preprint arXiv:1602.04676},
  year   = {2016}
}
R2 v1 2026-06-22T12:50:23.887Z