中文

通过高斯逼近实现$\tilde{\mathcal{O}}(1/N)$最优性间隙的游荡匪徒

最优化与控制 2025-05-27 v2 机器学习 概率论

摘要

我们研究具有NN个同质臂的有限时域游荡多臂匪徒(RMAB)问题。先前工作表明,当RMAB满足非退化条件时,基于流体逼近(捕捉系统均值动力学)的线性规划(LP)策略可实现指数小的最优性间隙。然而,RMAB通常是退化的,此时基于LP的策略可能导致每臂Θ(1/N)\Theta(1/\sqrt{N})的最优性间隙。本文提出了一种新颖的基于随机规划(SP)的策略,在唯一性假设下,对于退化RMAB实现了O~(1/N)\tilde{\mathcal{O}}(1/N)的最优性间隙。我们的方法基于构建一个高斯随机系统,该系统不仅捕捉RMAB动力学的均值,还捕捉其方差,从而比流体逼近更精确。然后我们求解该系统的随机规划以获得策略。这是首个为退化RMAB建立O~(1/N)\tilde{\mathcal{O}}(1/N)最优性间隙的结果。

关键词

引用

@article{arxiv.2410.15003,
  title  = {Achieving $\tilde{\mathcal{O}}(1/N)$ Optimality Gap in Restless Bandits through Gaussian Approximation},
  author = {Chen Yan and Weina Wang and Lei Ying},
  journal= {arXiv preprint arXiv:2410.15003},
  year   = {2025}
}

备注

53 pages, 5 figures. We weaken the assumption underlying our main theoretical result and also prove a more general version that holds without this condition