Grokking as Dimensional Phase Transition in Neural Networks
Abstract
Neural network grokking -- the abrupt memorization-to-generalization transition -- challenges our understanding of learning dynamics. Through finite-size scaling of gradient avalanche dynamics across eight model scales, we find that grokking is a \textit{dimensional phase transition}: effective dimensionality~ crosses from sub-diffusive (subcritical, ) to super-diffusive (supercritical, ) at generalization onset, exhibiting self-organized criticality (SOC). Crucially, reflects \textbf{gradient field geometry}, not network architecture: synthetic i.i.d.\ Gaussian gradients maintain regardless of graph topology, while real training exhibits dimensional excess from backpropagation correlations. The grokking-localized crossing -- robust across topologies -- offers new insight into the trainability of overparameterized networks.
Cite
@article{arxiv.2604.04655,
title = {Grokking as Dimensional Phase Transition in Neural Networks},
author = {Ping Wang},
journal= {arXiv preprint arXiv:2604.04655},
year = {2026}
}