English

Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms

Machine Learning 2025-10-31 v1 Machine Learning Statistics Theory Statistics Theory

Abstract

We consider a stochastic multi-armed bandit problem with i.i.d. rewards where the expected reward function is multimodal with at most m modes. We propose the first known computationally tractable algorithm for computing the solution to the Graves-Lai optimization problem, which in turn enables the implementation of asymptotically optimal algorithms for this bandit problem. The code for the proposed algorithms is publicly available at https://github.com/wilrev/MultimodalBandits

Keywords

Cite

@article{arxiv.2510.25811,
  title  = {Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms},
  author = {William Réveillard and Richard Combes},
  journal= {arXiv preprint arXiv:2510.25811},
  year   = {2025}
}

Comments

31 pages; NeurIPS 2025

R2 v1 2026-07-01T07:12:34.240Z