English

Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks

Machine Learning 2026-05-13 v1 Disordered Systems and Neural Networks Statistical Mechanics Information Theory Machine Learning math.IT

Abstract

We study the information-theoretic limits of learning a one-hidden-layer teacher network with hierarchical features from noisy queries, in the context of knowledge transfer to a smaller student model. We work in the high-dimensional regime where the teacher width kk scales linearly with the input dimension dd -- a setting that captures large-but-finite-width networks and has only recently become analytically tractable. Using a heuristic leave-one-out decoupling argument, validated numerically throughout, we derive asymptotically sharp characterizations of the Bayes-optimal generalization error and individual feature overlaps via a system of closed fixed-point equations. These equations reveal that feature learnability is governed by a sequence of sharp phase transitions: as data grows, teacher features become recoverable sequentially, each through a discontinuous jump in overlap. This sequential acquisition underlies a precise notion of \textit{effective width} kck_c -- the number of learnable features at a given data budget nn -- which unifies two distinct scaling regimes: a feature-learning regime in which the Bayes-optimal generalization error εBO\varepsilon^{\rm BO} scales as n1/(2β)1 n^{1/(2\beta)-1}, and a refinement regime in which it scales as n1n^{-1}, where β>1/2\beta>1/2 is the exponent of the power-law feature hierarchy. Both laws collapse to the single relation εBO=Θ(kcd/n)\varepsilon^{\rm BO}=\Theta(k_c d/n). We further show empirically that a student trained with \textsc{Adam} near the effective width kck_c achieves these optimal scaling laws (up to a small algorithmic gap), and provide an information-theoretic account of the associated scaling in model size.

Keywords

Cite

@article{arxiv.2605.10395,
  title  = {Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks},
  author = {Minh-Toan Nguyen and Jean Barbier},
  journal= {arXiv preprint arXiv:2605.10395},
  year   = {2026}
}