English

Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards

Machine Learning 2024-04-15 v4 Machine Learning

Abstract

We study KK-armed bandit problems where the reward distributions of the arms are all supported on the [0,1][0,1] interval. It has been a challenge to design regret-efficient randomized exploration algorithms in this setting. Maillard sampling \cite{maillard13apprentissage}, an attractive alternative to Thompson sampling, has recently been shown to achieve competitive regret guarantees in the sub-Gaussian reward setting \cite{bian2022maillard} while maintaining closed-form action probabilities, which is useful for offline policy evaluation. In this work, we propose the Kullback-Leibler Maillard Sampling (KL-MS) algorithm, a natural extension of Maillard sampling for achieving KL-style gap-dependent regret bound. We show that KL-MS enjoys the asymptotic optimality when the rewards are Bernoulli and has a worst-case regret bound of the form O(μ(1μ)KTlnK+KlnT)O(\sqrt{\mu^*(1-\mu^*) K T \ln K} + K \ln T), where μ\mu^* is the expected reward of the optimal arm, and TT is the time horizon length.

Keywords

Cite

@article{arxiv.2304.14989,
  title  = {Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards},
  author = {Hao Qin and Kwang-Sung Jun and Chicheng Zhang},
  journal= {arXiv preprint arXiv:2304.14989},
  year   = {2024}
}

Comments

Accepted by NeurIPS 2023

R2 v1 2026-06-28T10:20:59.285Z