English

Unifying Grokking and Double Descent

Machine Learning 2023-03-14 v1 Artificial Intelligence

Abstract

A principled understanding of generalization in deep learning may require unifying disparate observations under a single conceptual framework. Previous work has studied \emph{grokking}, a training dynamic in which a sustained period of near-perfect training performance and near-chance test performance is eventually followed by generalization, as well as the superficially similar \emph{double descent}. These topics have so far been studied in isolation. We hypothesize that grokking and double descent can be understood as instances of the same learning dynamics within a framework of pattern learning speeds. We propose that this framework also applies when varying model capacity instead of optimization steps, and provide the first demonstration of model-wise grokking.

Keywords

Cite

@article{arxiv.2303.06173,
  title  = {Unifying Grokking and Double Descent},
  author = {Xander Davies and Lauro Langosco and David Krueger},
  journal= {arXiv preprint arXiv:2303.06173},
  year   = {2023}
}

Comments

ML Safety Workshop, 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

R2 v1 2026-06-28T09:11:42.830Z