A scaling law of contextual persistence in human language
Abstract
Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d (, mean ; median ). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.
Keywords
Cite
@article{arxiv.2607.25184,
title = {A scaling law of contextual persistence in human language},
author = {Elan Barenholtz},
journal= {arXiv preprint arXiv:2607.25184},
year = {2026}
}
Comments
21 pages, 5 figures (plus 1 supplementary figure); Supplementary Information included