English

Convergence of gradient descent for learning linear neural networks

Machine Learning 2021-11-25 v2 Optimization and Control

Abstract

We study the convergence properties of gradient descent for training deep linear neural networks, i.e., deep matrix factorizations, by extending a previous analysis for the related gradient flow. We show that under suitable conditions on the step sizes gradient descent converges to a critical point of the loss function, i.e., the square loss in this article. Furthermore, we demonstrate that for almost all initializations gradient descent converges to a global minimum in the case of two layers. In the case of three or more layers we show that gradient descent converges to a global minimum on the manifold matrices of some fixed rank, where the rank cannot be determined a priori.

Keywords

Cite

@article{arxiv.2108.02040,
  title  = {Convergence of gradient descent for learning linear neural networks},
  author = {Gabin Maxime Nguegnang and Holger Rauhut and Ulrich Terstiege},
  journal= {arXiv preprint arXiv:2108.02040},
  year   = {2021}
}

Comments

Minor changes

R2 v1 2026-06-24T04:49:29.626Z