English

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

Machine Learning 2025-04-02 v2 Disordered Systems and Neural Networks Statistical Mechanics Information Theory Machine Learning math.IT

Abstract

We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width kk and input dimension dd are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size nn scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters kd+kkd+k and of data nn are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate n/d2n/d^2, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.

Keywords

Cite

@article{arxiv.2501.18530,
  title  = {Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation},
  author = {Jean Barbier and Francesco Camilli and Minh-Toan Nguyen and Mauro Pastore and Rudy Skerk},
  journal= {arXiv preprint arXiv:2501.18530},
  year   = {2025}
}

Comments

v2: 9 pages + appendix, 10 figures, 3 tables; added discussion on Gaussian inner weights (Fig. 2, 5 + Appendix H); added discussion on algorithmic complexity of specialisation (Appendix I and figures therein)

R2 v1 2026-06-28T21:26:03.729Z