English

Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories

Machine Learning 2026-04-29 v1 Artificial Intelligence

Abstract

We show that replacing the rolling SVD of AdamW updates with a rolling SVD of loss gradients changes the diagnostic by 1-2 orders of magnitude. Performing SVD on the loss gradient instead of the AdamW update increases the measured perturbative coupling between SED directions and Linear Centroid Hypothesis (LCH) features from Rˉk3 \bar{R}_k \approx 3 --9×9\times to 100100--330×330\times across four single-task modular arithmetic operations, eliminating the apparent operation dependence in the original measurement. On a multitask transformer with a shared encoder, update-based SED gives Rˉk1 \bar{R}_k \leq 1 -- an apparent failure of the diagnostic -- while per-operation gradient-based SED recovers Rˉk=20 \bar{R}_k = 20 --45×45\times across all four operations. Gradient aggregation across competing tasks is the main obstruction; performing SVD on per-task gradients resolves it. A causal intervention shows that constraining attention updates to any rank-3 subspace (whether SED-derived or random) accelerates grokking by approximately 2.3×2.3\times across random seeds and operations, while removing the rank-3 component has negligible effect under proper gradient-projection methodology. The SED-LCH coupling is therefore a strong diagnostic of where feature formation concentrates in parameter space, but it is not a unique causal pathway: the natural full-rank AdamW attention update is highly rank-redundant under our hyperparameters.

Keywords

Cite

@article{arxiv.2604.25143,
  title  = {Gradient-Direction Sensitivity Reveals Linear-Centroid Coupling Hidden by Optimizer Trajectories},
  author = {Yongzhong Xu},
  journal= {arXiv preprint arXiv:2604.25143},
  year   = {2026}
}

Comments

15 pages, 5 figures

R2 v1 2026-07-01T12:38:22.699Z