English

Connections Between Mirror Descent, Thompson Sampling and the Information Ratio

Machine Learning 2019-05-29 v1 Machine Learning

Abstract

The information-theoretic analysis by Russo and Van Roy (2014) in combination with minimax duality has proved a powerful tool for the analysis of online learning algorithms in full and partial information settings. In most applications there is a tantalising similarity to the classical analysis based on mirror descent. We make a formal connection, showing that the information-theoretic bounds in most applications can be derived from existing techniques for online convex optimisation. Besides this, for kk-armed adversarial bandits we provide an efficient algorithm with regret that matches the best information-theoretic upper bound and improve best known regret guarantees for online linear optimisation on p\ell_p-balls and bandits with graph feedback.

Keywords

Cite

@article{arxiv.1905.11817,
  title  = {Connections Between Mirror Descent, Thompson Sampling and the Information Ratio},
  author = {Julian Zimmert and Tor Lattimore},
  journal= {arXiv preprint arXiv:1905.11817},
  year   = {2019}
}
R2 v1 2026-06-23T09:29:02.700Z