中文

针对非平稳线性多臂赌博问题的探查-提交策略:面向潜在动态

机器学习 2025-10-21 v1 系统与控制 系统与控制 最优化与控制 机器学习

摘要

我们研究了一个奖励取决于动作和潜在状态的非平稳多臂赌博问题,其中后者由未知线性动力学控制。关键在于,状态动力学也取决于动作,导致短期和长期奖励之间存在张力。我们提出一种针对有限时间范围 TT 的探查-提交算法。 During the exploration phase, random Rademacher actions enable estimation of the Markov parameters of the linear dynamics, which characterize the action-reward relationship. In the commit phase, the algorithm uses the estimated parameters to design an optimized action sequence for long-term reward. Our proposed algorithm achieves O~(T2/3)\tilde{\mathcal{O}}(T^{2/3}) regret. Our analysis handles two key challenges: learning from temporally correlated rewards, and designing action sequences with optimal long-term reward. We address the first challenge by providing near-optimal sample complexity and error bounds for system identification using bilinear rewards. We address the second challenge by proving an equivalence with indefinite quadratic optimization over a hypercube, a known NP-hard problem. We provide a sub-optimality guarantee for this problem, enabling our regret upper bound. Lastly, we propose a semidefinite relaxation with Goemans-Williamson rounding as a practical approach.

关键词

引用

@article{arxiv.2510.16208,
  title  = {Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics},
  author = {Sunmook Choi and Yahya Sattar and Yassir Jedra and Maryam Fazel and Sarah Dean},
  journal= {arXiv preprint arXiv:2510.16208},
  year   = {2025}
}