English

A Neural Scaling Law from the Dimension of the Data Manifold

Machine Learning 2020-04-24 v1 Machine Learning

Abstract

When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law LNαL \propto N^{-\alpha} in the number of network parameters NN. This empirical scaling law holds for a wide variety of data modalities, and may persist over many orders of magnitude. The scaling law can be explained if neural models are effectively just performing regression on a data manifold of intrinsic dimension dd. This simple theory predicts that the scaling exponents α4/d\alpha \approx 4/d for cross-entropy and mean-squared error losses. We confirm the theory by independently measuring the intrinsic dimension and the scaling exponents in a teacher/student framework, where we can study a variety of dd and α\alpha by dialing the properties of random teacher networks. We also test the theory with CNN image classifiers on several datasets and with GPT-type language models.

Keywords

Cite

@article{arxiv.2004.10802,
  title  = {A Neural Scaling Law from the Dimension of the Data Manifold},
  author = {Utkarsh Sharma and Jared Kaplan},
  journal= {arXiv preprint arXiv:2004.10802},
  year   = {2020}
}

Comments

16+12 pages, 11+11 figures