English
Related papers

Related papers: From Boltzmann to Zipf through Shannon and Jaynes

200 papers

The pioneering research of G. K. Zipf on the relationship between word frequency and other word features led to the formulation of various linguistic laws. The most popular is Zipf's law for word frequencies. Here we focus on two laws that…

Computation and Language · Computer Science 2020-09-24 Bernardino Casas , Antoni Hernández-Fernández , Neus Català , Ramon Ferrer-i-Cancho , Jaume Baixeries

Zipf's law, and power laws in general, have attracted and continue to attract considerable attention in a wide variety of disciplines - from astronomy to demographics to software structure to economics to linguistics to zoology, and even…

Physics and Society · Physics 2013-05-10 Matt Visser

The binary many-step Markov chain with the step-like memory function is considered as a model for the analysis of rank distributions of words in stochastic symbolic dynamical systems. We prove that the envelope curve for this distribution…

History and Philosophy of Physics · Physics 2007-05-23 K. E. Kechedzhy O. V. Usatenko , V. A. Yampol'skii

Zipf's law seems to be ubiquitous in human languages and appears to be a universal property of complex communicating systems. Following the early proposal made by Zipf concerning the presence of a tension between the efforts of speaker and…

Adaptation and Self-Organizing Systems · Physics 2015-05-19 Bernat Corominas-Murtra , Jordi Fortuny , Ricard V. Solé

We use the formalism of 'Maximum Principle of Shannon's Entropy' to derive the general power law distribution function, using what seems to be a reasonable physical assumption, namely, the demand of a constant mean "internal order"…

Statistical Mechanics · Physics 2007-05-23 Yaniv Dover

Understanding the innovation process, that is the underlying mechanisms through which novelties emerge, diffuse and trigger further novelties is undoubtedly of fundamental importance in many areas (biology, linguistics, social science and…

Physics and Society · Physics 2021-10-29 Giacomo Aletti , Irene Crimaldi

The prevailing maximum likelihood estimators for inferring power law models from rank-frequency data are biased. The source of this bias is an inappropriate likelihood function. The correct likelihood function is derived and shown to be…

Applications · Statistics 2021-07-27 Charlie Pilgrim , Thomas T Hills

In this paper we combine statistical analysis of large text databases and simple stochastic models to explain the appearance of scaling laws in the statistics of word frequencies. Besides the sublinear scaling of the vocabulary size with…

Physics and Society · Physics 2014-11-05 Martin Gerlach , Eduardo G. Altmann

Many features from texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper we quantify how topological properties of word co-occurrence networks and…

Physics and Society · Physics 2013-02-20 Diego R. Amancio , Eduardo G. Altmann , Osvaldo N. Oliveira , Luciano da F. Costa

We analyse correspondence of a text to a simple probabilistic model. The model assumes that the words are selected independently from an infinite dictionary. The probability distribution correspond to the Zipf---Mandelbrot law. We count…

Tokenization is a fundamental step in natural language processing (NLP) and other sequence modeling domains, where the choice of vocabulary size significantly impacts model performance. Despite its importance, selecting an optimal…

Machine Learning · Computer Science 2025-07-31 Yanjin He , Qingkai Zeng , Meng Jiang

Human language, as a typical complex system, its organization and evolution is an attractive topic for both physical and cultural researchers. In this paper, we present the first exhaustive analysis of the text organization of human speech.…

Computation and Language · Computer Science 2015-01-08 Ruokuang Lin , Qianli D. Y. Ma , Chunhua Bian

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

Computation and Language · Computer Science 2017-10-04 Xiaoyong Yan , Petter Minnhagen

The inverse relationship between the length of a word and the frequency of its use, first identified by G.K. Zipf in 1935, is a classic empirical law that holds across a wide range of human languages. We demonstrate that length is one…

Computation and Language · Computer Science 2017-06-02 Stephan C. Meylan , Thomas L. Griffiths

Bayesian modelling and statistical text analysis rely on informed probability priors to encourage good solutions. This paper empirically analyses whether text in medical discharge reports follow Zipf's law, a commonly assumed statistical…

Computation and Language · Computer Science 2020-11-09 Juan C Quiroz , Liliana Laranjo , Catalin Tufanaru , Ahmet Baki Kocaballi , Dana Rezazadegan , Shlomo Berkovsky , Enrico Coiera

We demonstrate that large texts, representing human (English, Russian, Ukrainian) and artificial (C++, Java) languages, display quantitative patterns characterized by the Benford-like and Zipf laws. The frequency of a word following the…

Computation and Language · Computer Science 2018-03-13 Evgeny Shulzinger , Irina Legchenkova , Edward Bormashenko

Long-range correlations are found in symbolic sequences from human language, music and DNA. Determining the span of correlations in dolphin whistle sequences is crucial for shedding light on their communicative complexity. Dolphin whistles…

Neurons and Cognition · Quantitative Biology 2014-12-03 Ramon Ferrer-i-Cancho , Brenda McCowan

Zipf's law, which states that the probability of an observation is inversely proportional to its rank, has been observed in many domains. While there are models that explain Zipf's law in each of them, those explanations are typically…

Neurons and Cognition · Quantitative Biology 2016-07-06 Laurence Aitchison , Nicola Corradi , Peter E. Latham

The problem addressed concerns the determination of the average number of successive attempts of guessing a word of a certain length consisting of letters with given probabilities of occurrence. Both first- and second-order approximations…

Information Theory · Computer Science 2015-06-19 Kerstin Andersson

We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas model in statistical…

Computation and Language · Computer Science 2016-04-22 Weibing Deng , Armen E. Allahverdyan