English

Finite-time Regret Bound of a Bandit Algorithm for the Semi-bounded Support Model

Statistics Theory 2012-02-20 v2 Probability Statistics Theory

Abstract

In this paper we consider stochastic multiarmed bandit problems. Recently a policy, DMED, is proposed and proved to achieve the asymptotic bound for the model that each reward distribution is supported in a known bounded interval, e.g. [0,1]. However, the derived regret bound is described in an asymptotic form and the performance in finite time has been unknown. We inspect this policy and derive a finite-time regret bound by refining large deviation probabilities to a simple finite form. Further, this observation reveals that the assumption on the lower-boundedness of the support is not essential and can be replaced with a weaker one, the existence of the moment generating function.

Keywords

Cite

@article{arxiv.1202.2277,
  title  = {Finite-time Regret Bound of a Bandit Algorithm for the Semi-bounded Support Model},
  author = {Junya Honda and Akimichi Takemura},
  journal= {arXiv preprint arXiv:1202.2277},
  year   = {2012}
}
R2 v1 2026-06-21T20:17:43.299Z