追逐移动子空间:超越平稳性的低秩多臂老腰
摘要
许多多臂老腰部署(推荐、临床剂量、广告定向)共享两个事实:前人工作仅在孤立处理:奖励落在低维潜子空间内,且该子空间漂移。 stationary low-rank bandits 利用秩但在子空间变化时失效;non-stationary linear bandits 适应漂移但以 ambient rate 付出代价。我们研究 piecewise-stationary low-rank linear contextual bandits with scalar feedback:,其中秩为 的因子 在每个未知 segment 中保持恒定且可在边界处迁移。我们的结果在三个轴向上都紧凑。(i)识别边界。仅凭单次 scalar rewards,移动子空间可通过 rewards 的二次函数完成恢复,当且仅当满足三个探测侧条件:已知噪声方差、受控 state-noise 耦合、完整维探测支持。每个条件在 unrestricted-second-moment 问题中都是必要的,联合起来充要,这界定了可解区域的边界。(ii)算法与动态 regret。SPSC 在 isotropic probes 与 windowed projected ridge-UCB 利用内学习的 维子空间中交织;一种 CUSUM 变体在线发现 segment 边界。 costed dynamic regret 为 ,用内在秩取代 ambient 率。(iii)实证。基于 11 个基准(合成、UCI/MovieLens、半合成临床、ZOZOTOWN 生产日志数据),SPSC 在 时超过 non-stationary 和 low-rank 基线,匹配分析中的交叉点。据我们了解,这是首次刻画识别边界并实现该设置下内在秩 dynamic-regret 率的工作。
引用
@article{arxiv.2605.20269,
title = {Catching a Moving Subspace: Low-Rank Bandits Beyond Stationarity},
author = {Hamed Khosravi and Xiaoming Huo},
journal= {arXiv preprint arXiv:2605.20269},
year = {2026}
}