English

Linear Frequency Principle Model to Understand the Absence of Overfitting in Neural Networks

Machine Learning 2021-05-26 v1 Data Analysis, Statistics and Probability

Abstract

Why heavily parameterized neural networks (NNs) do not overfit the data is an important long standing open question. We propose a phenomenological model of the NN training to explain this non-overfitting puzzle. Our linear frequency principle (LFP) model accounts for a key dynamical feature of NNs: they learn low frequencies first, irrespective of microscopic details. Theory based on our LFP model shows that low frequency dominance of target functions is the key condition for the non-overfitting of NNs and is verified by experiments. Furthermore, through an ideal two-layer NN, we unravel how detailed microscopic NN training dynamics statistically gives rise to a LFP model with quantitative prediction power.

Keywords

Cite

@article{arxiv.2102.00200,
  title  = {Linear Frequency Principle Model to Understand the Absence of Overfitting in Neural Networks},
  author = {Yaoyu Zhang and Tao Luo and Zheng Ma and Zhi-Qin John Xu},
  journal= {arXiv preprint arXiv:2102.00200},
  year   = {2021}
}

Comments

to appear in Chinese Physics Letters