Logistic Bandit 中的近最优纯探索
机器学习
2025-02-10 v2 机器学习
摘要
Bandit 算法因其在现实场景中的实际应用而备受关注。然而,除了多臂或线性 bandit 等简单设置外,最优算法仍然稀缺。值得注意的是,对于广义线性模型(GLM)bandit 背景下的纯探索问题,目前尚不存在最优解。在本文中,我们缩小了这一差距,并为 Logistic Bandit 下的一般纯探索问题开发了首个 track-and-stop 算法,称为 Logistic Track-and-Stop(Log-TS)。Log-TS 是一种高效算法,在渐近意义上以对数因子的精度匹配了预期样本复杂度的实例特定下界的近似。
引用
@article{arxiv.2410.20640,
title = {Near Optimal Pure Exploration in Logistic Bandits},
author = {Eduardo Ochoa Rivera and Ambuj Tewari},
journal= {arXiv preprint arXiv:2410.20640},
year = {2025}
}
备注
25 pages, 2 figures. arXiv admin note: text overlap with arXiv:2006.16073 by other authors