Effects of momentum scaling for SGD
Optimization and Control
2022-10-24 v1
Abstract
The paper studies the properties of stochastic gradient methods with preconditioning. We focus on momentum updated preconditioners with momentum coefficient . Seeking to explain practical efficiency of scaled methods, we provide convergence analysis in a norm associated with preconditioner, and demonstrate that scaling allows one to get rid of gradients Lipschitz constant in convergence rates. Along the way, we emphasize important role of , undeservedly set to constant at the arbitrariness of various authors. Finally, we propose the explicit constructive formulas for adaptive and step size values.
Keywords
Cite
@article{arxiv.2210.11869,
title = {Effects of momentum scaling for SGD},
author = {Dmitry A. Pasechnyuk and Alexander Gasnikov and Martin Takáč},
journal= {arXiv preprint arXiv:2210.11869},
year = {2022}
}
Comments
19 pages, 14 figures