English

Generalization in Deep Networks: The Role of Distance from Initialization

Machine Learning 2019-01-15 v2 Artificial Intelligence Machine Learning

Abstract

Why does training deep neural networks using stochastic gradient descent (SGD) result in a generalization error that does not worsen with the number of parameters in the network? To answer this question, we advocate a notion of effective model capacity that is dependent on {\em a given random initialization of the network} and not just the training algorithm and the data distribution. We provide empirical evidences that demonstrate that the model capacity of SGD-trained deep networks is in fact restricted through implicit regularization of {\em the 2\ell_2 distance from the initialization}. We also provide theoretical arguments that further highlight the need for initialization-dependent notions of model capacity. We leave as open questions how and why distance from initialization is regularized, and whether it is sufficient to explain generalization.

Keywords

Cite

@article{arxiv.1901.01672,
  title  = {Generalization in Deep Networks: The Role of Distance from Initialization},
  author = {Vaishnavh Nagarajan and J. Zico Kolter},
  journal= {arXiv preprint arXiv:1901.01672},
  year   = {2019}
}

Comments

Spotlight paper at NeurIPS 2017 workshop on Deep Learning: Bridging Theory and Practice