English

An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits

Machine Learning 2016-05-30 v1

Abstract

We present an algorithm that achieves almost optimal pseudo-regret bounds against adversarial and stochastic bandits. Against adversarial bandits the pseudo-regret is O(Knlogn)O(K\sqrt{n \log n}) and against stochastic bandits the pseudo-regret is O(i(logn)/Δi)O(\sum_i (\log n)/\Delta_i). We also show that no algorithm with O(logn)O(\log n) pseudo-regret against stochastic bandits can achieve O~(n)\tilde{O}(\sqrt{n}) expected regret against adaptive adversarial bandits. This complements previous results of Bubeck and Slivkins (2012) that show O~(n)\tilde{O}(\sqrt{n}) expected adversarial regret with O((logn)2)O((\log n)^2) stochastic pseudo-regret.

Keywords

Cite

@article{arxiv.1605.08722,
  title  = {An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits},
  author = {Peter Auer and Chao-Kai Chiang},
  journal= {arXiv preprint arXiv:1605.08722},
  year   = {2016}
}
R2 v1 2026-06-22T14:11:28.111Z