English

The Batch Complexity of Bandit Pure Exploration

Machine Learning 2025-02-04 v1 Machine Learning

Abstract

In a fixed-confidence pure exploration problem in stochastic multi-armed bandits, an algorithm iteratively samples arms and should stop as early as possible and return the correct answer to a query about the arms distributions. We are interested in batched methods, which change their sampling behaviour only a few times, between batches of observations. We give an instance-dependent lower bound on the number of batches used by any sample efficient algorithm for any pure exploration task. We then give a general batched algorithm and prove upper bounds on its expected sample complexity and batch complexity. We illustrate both lower and upper bounds on best-arm identification and thresholding bandits.

Keywords

Cite

@article{arxiv.2502.01425,
  title  = {The Batch Complexity of Bandit Pure Exploration},
  author = {Adrienne Tuynman and Rémy Degenne},
  journal= {arXiv preprint arXiv:2502.01425},
  year   = {2025}
}
R2 v1 2026-06-28T21:30:42.672Z