English

Training behavior of deep neural network in frequency domain

Machine Learning 2019-11-04 v6 Artificial Intelligence Information Theory math.IT Statistics Theory Machine Learning Statistics Theory

Abstract

Why deep neural networks (DNNs) capable of overfitting often generalize well in practice is a mystery [#zhang2016understanding]. To find a potential mechanism, we focus on the study of implicit biases underlying the training process of DNNs. In this work, for both real and synthetic datasets, we empirically find that a DNN with common settings first quickly captures the dominant low-frequency components, and then relatively slowly captures the high-frequency ones. We call this phenomenon Frequency Principle (F-Principle). The F-Principle can be observed over DNNs of various structures, activation functions, and training algorithms in our experiments. We also illustrate how the F-Principle help understand the effect of early-stopping as well as the generalization of DNNs. This F-Principle potentially provides insights into a general principle underlying DNN optimization and generalization.

Keywords

Cite

@article{arxiv.1807.01251,
  title  = {Training behavior of deep neural network in frequency domain},
  author = {Zhi-Qin John Xu and Yaoyu Zhang and Yanyang Xiao},
  journal= {arXiv preprint arXiv:1807.01251},
  year   = {2019}
}

Comments

To appear in 2019 26th-International conference of neural information processing (ICONIP)