English

From SGD to Spectra: A Theory of Neural Network Weight Dynamics

Machine Learning 2026-02-10 v2

Abstract

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the microscopic dynamics of SGD to the macroscopic evolution of singular-value spectra in weight matrices. We derive exact SDEs showing that squared singular values follow Dyson Brownian motion with eigenvalue repulsion, and characterize stationary distributions as gamma-type densities with power-law tails, providing the first theoretical explanation for the empirically observed 'bulk+tail' spectral structure in trained networks. Through controlled experiments on transformer and MLP architectures, we validate our theoretical predictions and demonstrate quantitative agreement between SDE-based forecasts and observed spectral evolution, providing a rigorous foundation for understanding why deep learning works.

Keywords

Cite

@article{arxiv.2507.12709,
  title  = {From SGD to Spectra: A Theory of Neural Network Weight Dynamics},
  author = {Brian Richard Olsen and Sam Fatehmanesh and Frank Xiao and Adarsh Kumarappan and Anirudh Gajula},
  journal= {arXiv preprint arXiv:2507.12709},
  year   = {2026}
}