English

Bayesian Algorithms for Adversarial Online Learning: from Finite to Infinite Action Spaces

Machine Learning 2025-09-23 v5 Computer Science and Game Theory Statistics Theory Machine Learning Statistics Theory

Abstract

We develop a form Thompson sampling for online learning under full feedback - also known as prediction with expert advice - where the learner's prior is defined over the space of an adversary's future actions, rather than the space of experts. We show regret decomposes into regret the learner expected a priori, plus a prior-robustness-type term we call excess regret. In the classical finite-expert setting, this recovers optimal rates. As an initial step towards practical online learning in settings with a potentially-uncountably-infinite number of experts, we show that Thompson sampling over the dd-dimensional unit cube, using a certain Gaussian process prior widely-used in the Bayesian optimization literature, has a O(βTdlog(1+dλβ))\mathcal{O}\Big(\beta\sqrt{Td\log(1+\sqrt{d}\frac{\lambda}{\beta})}\Big) rate against a β\beta-bounded λ\lambda-Lipschitz adversary.

Keywords

Cite

@article{arxiv.2502.14790,
  title  = {Bayesian Algorithms for Adversarial Online Learning: from Finite to Infinite Action Spaces},
  author = {Alexander Terenin and Jeffrey Negrea},
  journal= {arXiv preprint arXiv:2502.14790},
  year   = {2025}
}

Comments

This version renames the paper: its original name was "An Adversarial Analysis of Thompson Sampling for Full-information Online Learning: from Finite to Infinite Action Spaces"

R2 v1 2026-06-28T21:51:43.844Z