English

Scaling limit of the Random Language Model

Disordered Systems and Neural Networks 2026-06-26 v1 Statistical Mechanics Computation and Language

Abstract

We develop a quantitative theory of the Random Language Model (RLM), an ensemble of stochastic context-free grammars, in a scaling limit where the number of hidden symbols NN \to \infty while the grammar temperature ϵ~d0\tilde{\epsilon}_d \to 0 at fixed x=ϵ~dlogNx = {\tilde\epsilon}_d \log N. In this limit, the model admits a controlled description based on a large-deviation principle over rule-usage patterns. A semi-annealed approximation maps the problem to a class of Random Energy Models with nontrivial combinatorics. We show that the RLM exhibits a condensation transition at a critical value xc=1/8x_c=1/8, below which rule usage concentrates and language statistics acquire a nontrivial dependence on corpus length. A second characteristic scale at x=1/2x=1/2 marks the onset of entropy reduction from its maximal value. Across these regimes, we derive explicit scaling laws for the number of distinct rules, entropy, and related observables, identifying distinct scaling, saturation, and critical regimes controlled by the interplay of grammar size, corpus length, and temperature. The theory resolves previous ambiguities regarding the existence of a thermodynamic transition and explains the slow approach to the large-NN limit as a consequence of the dependence on logN\log N. It further provides a unified framework in which universal statistical properties of language emerge from typical realizations of generative grammars, with implications for both natural language statistics and the behavior of large language models.

Cite

@article{arxiv.2606.28105,
  title  = {Scaling limit of the Random Language Model},
  author = {Eric De Giuli},
  journal= {arXiv preprint arXiv:2606.28105},
  year   = {2026}
}

Comments

17 pages + 14 pages SI

R2 v1 2026-07-22T20:11:37.557Z