English

Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee

Machine Learning 2025-08-15 v1

Abstract

Online restless multi-armed bandits (RMABs) typically assume that each arm follows a stationary Markov Decision Process (MDP) with fixed state transitions and rewards. However, in real-world applications like healthcare and recommendation systems, these assumptions often break due to non-stationary dynamics, posing significant challenges for traditional RMAB algorithms. In this work, we specifically consider NN-armd RMAB with non-stationary transition constrained by bounded variation budgets BB. Our proposed \rmab\; algorithm integrates sliding window reinforcement learning (RL) with an upper confidence bound (UCB) mechanism to simultaneously learn transition dynamics and their variations. We further establish that \rmab\; achieves O~(N2B14T34)\widetilde{\mathcal{O}}(N^2 B^{\frac{1}{4}} T^{\frac{3}{4}}) regret bound by leveraging a relaxed definition of regret, providing a foundational theoretical framework for non-stationary RMAB problems for the first time.

Keywords

Cite

@article{arxiv.2508.10804,
  title  = {Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee},
  author = {Yu-Heng Hung and Ping-Chun Hsieh and Kai Wang},
  journal= {arXiv preprint arXiv:2508.10804},
  year   = {2025}
}