分布式带动量随机梯度下降法最后迭代收敛性分析
最优化与控制
2025-05-19 v1
摘要
分布式随机梯度方法广泛用于保留数据隐私并确保大规模学习任务的可扩展性。虽然现有理论针对分布式动量随机梯度下降(mSGD)主要关注时间平均收敛,但更实用的最后迭代收敛性仍受到忽视。在本工作中,我们在非凸设置下分析了分布式mSGD的最后迭代收敛行为,基于经典的Robbins-Monro步长方案。我们证明了最后迭代的几乎sure收敛和收敛,并推导了收敛率。我们进一步表明,动量可以加速早期阶段的收敛,并提供实验支持我们的理论。
引用
@article{arxiv.2505.10889,
title = {Convergence Analysis of the Last Iterate in Distributed Stochastic Gradient Descent with Momentum},
author = {Difei Cheng and Ruinan Jin and Hong Qiao and Bo Zhang},
journal= {arXiv preprint arXiv:2505.10889},
year = {2025}
}
备注
16 pages