English

Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set

Computation and Language 2025-05-20 v3

Abstract

Auto-regressive language models (LMs) have been widely used to generate data in data-scarce domains to train new LMs, compensating for the scarcity of real-world data. Previous work experimentally found that LMs collapse when trained on recursively generated data. This paper presents a theoretical proof: once a corpus (such as a subset of the World Wide Web) begins to incorporate generated data and no new real-world data is added to the corpus, then no matter how small the amount of data each LM generates and contributes to the corpus, LM collapse is inevitable after sufficient time. This finding suggests that attempts to mitigate collapse by limiting the quantity of synthetic data in the corpus are fundamentally insufficient. Instead, avoiding collapse hinges on ensuring the quality of synthetic data.

Keywords

Cite

@article{arxiv.2412.14872,
  title  = {Theoretical Proof that Auto-regressive Language Models Collapse when Real-world Data is a Finite Set},
  author = {Lecheng Wang and Xianjie Shi and Ge Li and Jia Li and Xuanming Zhang and Yihong Dong and Wenpin Jiao and Hong Mei},
  journal= {arXiv preprint arXiv:2412.14872},
  year   = {2025}
}

Comments

20 pages, 3 figures

R2 v1 2026-06-28T20:42:16.171Z