English

Efficient Online Bandit Multiclass Learning with $\tilde{O}(\sqrt{T})$ Regret

Machine Learning 2018-01-19 v3 Machine Learning

Abstract

We present an efficient second-order algorithm with O~(1ηT)\tilde{O}(\frac{1}{\eta}\sqrt{T}) regret for the bandit online multiclass problem. The regret bound holds simultaneously with respect to a family of loss functions parameterized by η\eta, for a range of η\eta restricted by the norm of the competitor. The family of loss functions ranges from hinge loss (η=0\eta=0) to squared hinge loss (η=1\eta=1). This provides a solution to the open problem of (J. Abernethy and A. Rakhlin. An efficient bandit algorithm for T\sqrt{T}-regret in online multiclass prediction? In COLT, 2009). We test our algorithm experimentally, showing that it also performs favorably against earlier algorithms.

Keywords

Cite

@article{arxiv.1702.07958,
  title  = {Efficient Online Bandit Multiclass Learning with $\tilde{O}(\sqrt{T})$ Regret},
  author = {Alina Beygelzimer and Francesco Orabona and Chicheng Zhang},
  journal= {arXiv preprint arXiv:1702.07958},
  year   = {2018}
}

Comments

22 pages, 2 figures; ICML 2017; this version includes additional discussions of Newtron, and a variant of SOBA that directly uses an online exp-concave optimization oracle

R2 v1 2026-06-22T18:28:32.068Z