English

Grokking as Dimensional Phase Transition in Neural Networks

Machine Learning 2026-04-07 v1 Disordered Systems and Neural Networks Artificial Intelligence Adaptation and Self-Organizing Systems

Abstract

Neural network grokking -- the abrupt memorization-to-generalization transition -- challenges our understanding of learning dynamics. Through finite-size scaling of gradient avalanche dynamics across eight model scales, we find that grokking is a \textit{dimensional phase transition}: effective dimensionality~DD crosses from sub-diffusive (subcritical, D<1D < 1) to super-diffusive (supercritical, D>1D > 1) at generalization onset, exhibiting self-organized criticality (SOC). Crucially, DD reflects \textbf{gradient field geometry}, not network architecture: synthetic i.i.d.\ Gaussian gradients maintain D1D \approx 1 regardless of graph topology, while real training exhibits dimensional excess from backpropagation correlations. The grokking-localized D(t)D(t) crossing -- robust across topologies -- offers new insight into the trainability of overparameterized networks.

Keywords

Cite

@article{arxiv.2604.04655,
  title  = {Grokking as Dimensional Phase Transition in Neural Networks},
  author = {Ping Wang},
  journal= {arXiv preprint arXiv:2604.04655},
  year   = {2026}
}
R2 v1 2026-07-01T11:55:17.505Z