任意学习率下的随机梯度 bandit 的全局收敛性
机器学习
2025-02-12 v1
摘要
我们通过揭示随机梯度 bandit 算法使用任意常数学习率即可几乎必然收敛至全局最优策略,来对该算法提供新的理解。该结果表明,即使在标准光滑性和噪声控制假设失效的情况下,随机梯度算法仍能适当地平衡探索与开发。证明基于对动作采样率和累积进度与噪声之间新颖关系的发现,并拓展了当前对简单随机梯度方法在 bandit 场景中行为的理解。
引用
@article{arxiv.2502.07141,
title = {Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates},
author = {Jincheng Mei and Bo Dai and Alekh Agarwal and Sharan Vaswani and Anant Raj and Csaba Szepesvari and Dale Schuurmans},
journal= {arXiv preprint arXiv:2502.07141},
year = {2025}
}
备注
Updated version for a paper published at NeurIPS 2024