Generalization in Deep Networks: The Role of Distance from Initialization
Abstract
Why does training deep neural networks using stochastic gradient descent (SGD) result in a generalization error that does not worsen with the number of parameters in the network? To answer this question, we advocate a notion of effective model capacity that is dependent on {\em a given random initialization of the network} and not just the training algorithm and the data distribution. We provide empirical evidences that demonstrate that the model capacity of SGD-trained deep networks is in fact restricted through implicit regularization of {\em the distance from the initialization}. We also provide theoretical arguments that further highlight the need for initialization-dependent notions of model capacity. We leave as open questions how and why distance from initialization is regularized, and whether it is sufficient to explain generalization.
Keywords
Cite
@article{arxiv.1901.01672,
title = {Generalization in Deep Networks: The Role of Distance from Initialization},
author = {Vaishnavh Nagarajan and J. Zico Kolter},
journal= {arXiv preprint arXiv:1901.01672},
year = {2019}
}
Comments
Spotlight paper at NeurIPS 2017 workshop on Deep Learning: Bridging Theory and Practice