English

Regret Minimization in Heavy-Tailed Bandits

Machine Learning 2021-02-09 v1 Machine Learning

Abstract

We revisit the classic regret-minimization problem in the stochastic multi-armed bandit setting when the arm-distributions are allowed to be heavy-tailed. Regret minimization has been well studied in simpler settings of either bounded support reward distributions or distributions that belong to a single parameter exponential family. We work under the much weaker assumption that the moments of order (1+ϵ)(1+\epsilon) are uniformly bounded by a known constant B, for some given ϵ>0\epsilon > 0. We propose an optimal algorithm that matches the lower bound exactly in the first-order term. We also give a finite-time bound on its regret. We show that our index concentrates faster than the well known truncated or trimmed empirical mean estimators for the mean of heavy-tailed distributions. Computing our index can be computationally demanding. To address this, we develop a batch-based algorithm that is optimal up to a multiplicative constant depending on the batch size. We hence provide a controlled trade-off between statistical optimality and computational cost.

Keywords

Cite

@article{arxiv.2102.03734,
  title  = {Regret Minimization in Heavy-Tailed Bandits},
  author = {Shubhada Agrawal and Sandeep Juneja and Wouter M. Koolen},
  journal= {arXiv preprint arXiv:2102.03734},
  year   = {2021}
}

Comments

35 pages, 2 figures

R2 v1 2026-06-23T22:54:34.658Z