English

Small nonlinearities in activation functions create bad local minima in neural networks

Machine Learning 2019-05-29 v4 Optimization and Control Machine Learning

Abstract

We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with "slightest" nonlinearity, the empirical risks have spurious local minima in most cases. Our results thus indicate that in general "no spurious local minima" is a property limited to deep linear networks, and insights obtained from linear networks may not be robust. Specifically, for ReLU(-like) networks we constructively prove that for almost all practical datasets there exist infinitely many local minima. We also present a counterexample for more general activations (sigmoid, tanh, arctan, ReLU, etc.), for which there exists a bad local minimum. Our results make the least restrictive assumptions relative to existing results on spurious local optima in neural networks. We complete our discussion by presenting a comprehensive characterization of global optimality for deep linear networks, which unifies other results on this topic.

Keywords

Cite

@article{arxiv.1802.03487,
  title  = {Small nonlinearities in activation functions create bad local minima in neural networks},
  author = {Chulhee Yun and Suvrit Sra and Ali Jadbabaie},
  journal= {arXiv preprint arXiv:1802.03487},
  year   = {2019}
}

Comments

33 pages, appeared at ICLR 2019

R2 v1 2026-06-23T00:17:39.591Z