Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology
Machine Learning
2020-12-21 v3 Machine Learning
Abstract
Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with ( being the number of training samples). In this paper, we prove that, for deep networks, a single layer of width following the input layer suffices to ensure a similar guarantee. In particular, all the remaining layers are allowed to have constant widths, and form a pyramidal topology. We show an application of our result to the widely used LeCun's initialization and obtain an over-parameterization requirement for the single wide layer of order
Keywords
Cite
@article{arxiv.2002.07867,
title = {Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology},
author = {Quynh Nguyen and Marco Mondelli},
journal= {arXiv preprint arXiv:2002.07867},
year = {2020}
}
Comments
Accepted at NeurIPS 2020