Dimensional Criticality at Grokking Across MLPs and Transformers
Abstract
Abrupt transitions between distinct dynamical regimes are a hallmark of complex systems. Grokking in deep neural networks provides a striking example -- an abrupt transition from memorization to generalization long after training accuracy saturates -- yet robust macroscopic signatures of this transition remain elusive. Here we introduce \textbf{TDU--OFC} (Thresholded Diffusion Update--Olami-Feder-Christensen), an offline avalanche probe that converts gradient snapshots into cascade statistics and extracts a \emph{macroscopic observable} -- the time-resolved effective cascade dimension -- via grokking-aligned finite-size scaling. Across Transformers trained on modular addition and MLPs trained on XOR, we discover a localized dynamical crossing of the Gaussian diffusion baseline precisely at the generalization transition. The crossing direction is task-dependent: modular addition descends through (approaching from ), while XOR ascends (from ). This opposite-direction convergence is consistent with attraction toward a candidate shared critical manifold, rather than trivial residence near . Negative controls confirm this picture: ungrokked runs remain supercritical () and never enter the post-transition regime. In addition, avalanche distributions exhibit heavy tails and finite-size scaling consistent with the dimensional exponent extracted from . Shadow-probe controls () confirm that is non-invasive, and grokked trajectories diverge from ungrokked ones in some -- epochs before the behavioral transition.
Cite
@article{arxiv.2604.16431,
title = {Dimensional Criticality at Grokking Across MLPs and Transformers},
author = {Ping Wang},
journal= {arXiv preprint arXiv:2604.16431},
year = {2026}
}