English

Near-Optimal Regret in Adversarial Kernel Bandits

Machine Learning 2026-05-27 v1

Abstract

We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbert space (RKHS). We propose an exponential-weights algorithm built on a regularized importance-weighted loss estimator, together with an explicit correction term that cancels the bias introduced by the regularization. Our main result bounds the regret by O~(Td(λ)logX)\widetilde{{O}}\big(\sqrt{T\, d_*(\lambda)\,\log|{X}|}\big), where d(λ)d_*(\lambda) is a widely-adopted notion of effective dimension that captures the complexity of the kernel. Up to logarithmic factors, this matches the known rate achieved in the related stochastic kernel bandit problem. A notable application is the Mat\'ern(ν,d)(\nu,d) kernel with smoothness parameter ν\nu on Rd\mathbb{R}^d, for which our bound specializes to O~(T(ν+d)/(2ν+d))\widetilde{{O}}\big(T^{(\nu+d)/(2\nu+d)}\big), improving over the best-known prior rate of Chatterji et al. [2019] while simultaneously removing the rank-one adversary assumption required by their analysis. Moreover, this rate is the same as the known optimal rate for stochastic kernel bandits, and also matches a lower bound from concurrent work up to a logT\log T factor.

Keywords

Cite

@article{arxiv.2605.26585,
  title  = {Near-Optimal Regret in Adversarial Kernel Bandits},
  author = {Yu-Jie Zhang and Hao Qiu and Jonathan Scarlett and Kevin Jamieson},
  journal= {arXiv preprint arXiv:2605.26585},
  year   = {2026}
}
R2 v1 2026-07-22T07:33:52.243Z