English

A scaling law of contextual persistence in human language

Computation and Language 2026-07-28 v1

Abstract

Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d (P(d)dαP(d) \propto d^{-\alpha}, mean α=1.04\alpha = 1.04; median r2=0.96r^2 = 0.96). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.

Keywords

Cite

@article{arxiv.2607.25184,
  title  = {A scaling law of contextual persistence in human language},
  author = {Elan Barenholtz},
  journal= {arXiv preprint arXiv:2607.25184},
  year   = {2026}
}

Comments

21 pages, 5 figures (plus 1 supplementary figure); Supplementary Information included