English

Toy Models of Superposition

Machine Learning 2022-09-23 v1

Abstract

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a toy model where polysemanticity can be fully understood, arising as a result of models storing additional sparse features in "superposition." We demonstrate the existence of a phase change, a surprising connection to the geometry of uniform polytopes, and evidence of a link to adversarial examples. We also discuss potential implications for mechanistic interpretability.

Keywords

Cite

@article{arxiv.2209.10652,
  title  = {Toy Models of Superposition},
  author = {Nelson Elhage and Tristan Hume and Catherine Olsson and Nicholas Schiefer and Tom Henighan and Shauna Kravec and Zac Hatfield-Dodds and Robert Lasenby and Dawn Drain and Carol Chen and Roger Grosse and Sam McCandlish and Jared Kaplan and Dario Amodei and Martin Wattenberg and Christopher Olah},
  journal= {arXiv preprint arXiv:2209.10652},
  year   = {2022}
}

Comments

Also available at https://transformer-circuits.pub/2022/toy_model/index.html

R2 v1 2026-06-28T01:51:20.789Z