English

Multiscale Non-stationary Stochastic Bandits

Machine Learning 2020-02-14 v1 Machine Learning

Abstract

Classic contextual bandit algorithms for linear models, such as LinUCB, assume that the reward distribution for an arm is modeled by a stationary linear regression. When the linear regression model is non-stationary over time, the regret of LinUCB can scale linearly with time. In this paper, we propose a novel multiscale changepoint detection method for the non-stationary linear bandit problems, called Multiscale-LinUCB, which actively adapts to the changing environment. We also provide theoretical analysis of regret bound for Multiscale-LinUCB algorithm. Experimental results show that our proposed Multiscale-LinUCB algorithm outperforms other state-of-the-art algorithms in non-stationary contextual environments.

Keywords

Cite

@article{arxiv.2002.05289,
  title  = {Multiscale Non-stationary Stochastic Bandits},
  author = {Qin Ding and Cho-Jui Hsieh and James Sharpnack},
  journal= {arXiv preprint arXiv:2002.05289},
  year   = {2020}
}
R2 v1 2026-06-23T13:40:16.997Z