English

Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics

Machine Learning 2021-05-11 v1 Neural and Evolutionary Computing Probability

Abstract

Yang (2020a) recently showed that the Neural Tangent Kernel (NTK) at initialization has an infinite-width limit for a large class of architectures including modern staples such as ResNet and Transformers. However, their analysis does not apply to training. Here, we show the same neural networks (in the so-called NTK parametrization) during training follow a kernel gradient descent dynamics in function space, where the kernel is the infinite-width NTK. This completes the proof of the *architectural universality* of NTK behavior. To achieve this result, we apply the Tensor Programs technique: Write the entire SGD dynamics inside a Tensor Program and analyze it via the Master Theorem. To facilitate this proof, we develop a graphical notation for Tensor Programs.

Keywords

Cite

@article{arxiv.2105.03703,
  title  = {Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics},
  author = {Greg Yang and Etai Littwin},
  journal= {arXiv preprint arXiv:2105.03703},
  year   = {2021}
}

Comments

ICML 2021

R2 v1 2026-06-24T01:54:11.976Z