English

Adversarial Bandits against Arbitrary Strategies

Machine Learning 2025-11-26 v7 Machine Learning

Abstract

We study the adversarial bandit problem against arbitrary strategies, where the difficulty is captured by an unknown parameter SS, which is the number of switches in the best arm in hindsight. To handle this problem, we adopt the master-base framework using the online mirror descent method (OMD). We first provide a master-base algorithm with simple OMD, achieving O~(S1/2K1/3T2/3)\tilde{O}(S^{1/2}K^{1/3}T^{2/3}), in which T2/3T^{2/3} comes from the variance of loss estimators. To mitigate the impact of the variance, we propose using adaptive learning rates for OMD and achieve O~(min{SKTρ,SKT})\tilde{O}(\min\{\sqrt{SKT\rho},S\sqrt{KT}\}), where ρ\rho is a variance term for loss estimators.

Keywords

Cite

@article{arxiv.2205.14839,
  title  = {Adversarial Bandits against Arbitrary Strategies},
  author = {Jung-hun Kim and Se-Young Yun},
  journal= {arXiv preprint arXiv:2205.14839},
  year   = {2025}
}
R2 v1 2026-06-24T11:32:38.784Z