中文

基于新颖停止时间技术的AdaGrad非凸优化稳定性与收敛性分析

最优化与控制 2024-12-31 v3 机器学习 机器学习

摘要

自适应梯度优化器(AdaGrad)通过动态调整基于迭代梯度的学习率,已成为深度学习中的强大工具。这些自适应方法在 various deep learning tasks 中取得显著成功,超越随机梯度下降。然而,尽管AdaGrad作为自适应优化的基石,其理论分析尚未充分 addresses key aspects such as asymptotic convergence and non-asymptotic convergence rates in non-convex optimization scenarios。本研究旨在为AdaGrad提供全面分析,弥补文献中的现有空白。我们引入一种来自概率论的 new stopping time technique, which allows us to establish the stability of AdaGrad under mild conditions。我们进一步推导出AdaGrad的 almost sure 收敛和均方收敛。此外,我们展示了 AdaGrad 在 average-squared gradients in expectation 下具有 near-optimal non-asymptotic 收敛率,这强于 existing high-probability results。本 work developing 的技术可能对 future research on other adaptive stochastic algorithms 具有独立的兴趣。

关键词

引用

@article{arxiv.2409.05023,
  title  = {Stability and convergence analysis of AdaGrad for non-convex optimization via novel stopping time-based techniques},
  author = {Ruinan Jin and Xiaoyu Wang and Baoxiang Wang},
  journal= {arXiv preprint arXiv:2409.05023},
  year   = {2024}
}

备注

51 pages