中文

任意学习率下的随机梯度 bandit 的全局收敛性

机器学习 2025-02-12 v1

摘要

我们通过揭示随机梯度 bandit 算法使用任意常数学习率即可几乎必然收敛至全局最优策略,来对该算法提供新的理解。该结果表明,即使在标准光滑性和噪声控制假设失效的情况下,随机梯度算法仍能适当地平衡探索与开发。证明基于对动作采样率和累积进度与噪声之间新颖关系的发现,并拓展了当前对简单随机梯度方法在 bandit 场景中行为的理解。

关键词

引用

@article{arxiv.2502.07141,
  title  = {Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates},
  author = {Jincheng Mei and Bo Dai and Alekh Agarwal and Sharan Vaswani and Anant Raj and Csaba Szepesvari and Dale Schuurmans},
  journal= {arXiv preprint arXiv:2502.07141},
  year   = {2025}
}

备注

Updated version for a paper published at NeurIPS 2024