English

Best Arm Identification with Minimal Regret

Machine Learning 2024-09-30 v1 Information Theory math.IT Machine Learning

Abstract

Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret. This innovative variant of the multi-armed bandit problem elegantly amalgamates two of its most ubiquitous objectives: regret minimization and BAI. More precisely, the agent's goal is to identify the best arm with a prescribed confidence level δ\delta, while minimizing the cumulative regret up to the stopping time. Focusing on single-parameter exponential families of distributions, we leverage information-theoretic techniques to establish an instance-dependent lower bound on the expected cumulative regret. Moreover, we present an intriguing impossibility result that underscores the tension between cumulative regret and sample complexity in fixed-confidence BAI. Complementarily, we design and analyze the Double KL-UCB algorithm, which achieves asymptotic optimality as the confidence level tends to zero. Notably, this algorithm employs two distinct confidence bounds to guide arm selection in a randomized manner. Our findings elucidate a fresh perspective on the inherent connections between regret minimization and BAI.

Keywords

Cite

@article{arxiv.2409.18909,
  title  = {Best Arm Identification with Minimal Regret},
  author = {Junwen Yang and Vincent Y. F. Tan and Tianyuan Jin},
  journal= {arXiv preprint arXiv:2409.18909},
  year   = {2024}
}

Comments

Preprint

R2 v1 2026-06-28T18:59:46.334Z