中文

具有可证保证的非平稳无约束多臂老虎机

机器学习 2025-08-15 v1

摘要

在线无约束多臂老虎机 (RMAB) 通常假设每根臂遵循具有固定状态转移和奖励的平稳马尔可夫决策过程 (MDP)。然而,在医疗和推荐系统等实际应用中,这些假设常因非平稳动态而失效,这对传统 RMAB 算法构成重大挑战。本文特别考虑受限于 bounded variation 预算 BBNN 臂非平稳 RMAB。我们提出的 \text{rmab}\; 算法集成了滑动窗口强化学习 (RL) 与上置置信界 (UCB) 机制,同时学习状态转移动力学及其变化。我们进一步建立了 \text{rmab}\; 达到 O~(N2B14T34)\widetilde{\mathcal{O}}(N^2 B^{\frac{1}{4}} T^{\frac{3}{4}}) 的 regret 上界,通过松化的 regret 定义,首次为非平稳 RMAB 问题提供了基础性理论框架。

关键词

引用

@article{arxiv.2508.10804,
  title  = {Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee},
  author = {Yu-Heng Hung and Ping-Chun Hsieh and Kai Wang},
  journal= {arXiv preprint arXiv:2508.10804},
  year   = {2025}
}