English

Rank dynamics of word usage at multiple scales

Physics and Society 2026-02-04 v1

Abstract

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore whether word use is similar across languages, and if so, whether these generic features appear at different scales of language structure. Here we use the Google Books NN-grams dataset to analyze the temporal evolution of word usage in several languages. We apply measures proposed recently to study rank dynamics, such as the diversity of NN-grams in a given rank, the probability that an NN-gram changes rank between successive time intervals, the rank entropy, and the rank complexity. Using different methods, results show that there are generic properties for different languages at different scales, such as a core of words necessary to minimally understand a language. We also propose a null model to explore the relevance of linguistic structure across multiple scales, concluding that NN-gram statistics cannot be reduced to word statistics. We expect our results to be useful in improving text prediction algorithms, as well as in shedding light on the large-scale features of language use, beyond linguistic and cultural differences across human populations.

Keywords

Cite

@article{arxiv.1802.07258,
  title  = {Rank dynamics of word usage at multiple scales},
  author = {José A. Morales and Ewan Colman and Sergio Sánchez and Fernanda Sánchez-Puig and Carlos Pineda and Gerardo Iñiguez and Germinal Cocho and Jorge Flores and Carlos Gershenson},
  journal= {arXiv preprint arXiv:1802.07258},
  year   = {2026}
}

Comments

19 pages (main text) + 24 pages (supplementary information)