English

Phase Transitions for Feature Learning in Neural Networks

Machine Learning 2026-02-27 v2 Statistics Theory Statistics Theory

Abstract

According to a popular viewpoint, neural networks learn from data by first identifying low-dimensional representations, and subsequently fitting the best model in this space. Recent works provide a formalization of this phenomenon when learning multi-index models. In this setting, we are given nn i.i.d. pairs (xi,yi)({\boldsymbol x}_i,y_i), where the covariate vectors xiRd{\boldsymbol x}_i\in\mathbb{R}^d are isotropic, and responses yiy_i only depend on xi{\boldsymbol x}_i through a kk-dimensional projection ΘTxi{\boldsymbol \Theta}_*^{{\sf T}}{\boldsymbol x}_i. Feature learning amounts to learning the latent space spanned by Θ{\boldsymbol \Theta}_*. In this context, we study the gradient descent dynamics of two-layer neural networks under the proportional asymptotics n,dn,d\to\infty, n/dδn/d\to\delta, while the dimension of the latent space kk and the number of hidden neurons mm are kept fixed. Earlier work establishes that feature learning via polynomial-time algorithms is possible if δ>δalg\delta> \delta_{\text{alg}}, for δalg\delta_{\text{alg}} a threshold depending on the data distribution, and is impossible (within a certain class of algorithms) below δalg\delta_{\text{alg}}. Here we derive an analogous threshold δNN\delta_{\text{NN}} for two-layer networks. Our characterization of δNN\delta_{\text{NN}} opens the way to study the dependence of learning dynamics on the network architecture and training algorithm. The threshold δNN\delta_{\text{NN}} is determined by the following scenario. Training first visits points for which the gradient of the empirical risk is large and learns the directions spanned by these gradients. Then the gradient becomes smaller and the dynamics becomes dominated by negative directions of the Hessian. The threshold δNN\delta_{\text{NN}} corresponds to a phase transition in the spectrum of the Hessian in this second phase.

Keywords

Cite

@article{arxiv.2602.01434,
  title  = {Phase Transitions for Feature Learning in Neural Networks},
  author = {Andrea Montanari and Zihao Wang},
  journal= {arXiv preprint arXiv:2602.01434},
  year   = {2026}
}

Comments

75 pages; 17 pdf figures; v2 is a minor revision of v1

R2 v1 2026-07-01T09:30:33.667Z