English
Related papers

Related papers: From Boltzmann to Zipf through Shannon and Jaynes

200 papers

Biosemiosis is a process of choice-making between simultaneously alternative options. It is well-known that, when sufficiently young children encounter a new word, they tend to interpret it as pointing to a meaning that does not have a word…

Computation and Language · Computer Science 2022-09-22 David Carrera-Casado , Ramon Ferrer-i-Cancho

The Zipf distribution is a probability distribution widely used by scientists from various disciplines due to its ubiquity. Some of these areas include linguistics, physics, genetics, and sociology, among others. In this paper, it is proved…

Statistics Theory · Mathematics 2025-11-17 Marta Pérez-Casany , Ariel Duarte-López , Jordi Valero

In the Yule-Simon process, selection of words follows the preferential attachment mechanism, resulting in the power-law growth in the cumulative number of individual word occurrences. This is derived using mean-field approximation, assuming…

Statistical Mechanics · Physics 2016-05-04 Yasuhiro Hashimoto

Zipf's law implies the statistical distributions of hyperbolic type, which can describe the properties of stability and entropy loss in linguistics. We present the information theory from which follows that if the system is described by…

Biological Physics · Physics 2014-03-31 K. Lukierska-Walasek , K. Topolski , K. Trojanowski

Words are sequences of letters over a finite alphabet. We study two intimately related topics for this object: quasi-randomness and limit theory. With respect to the first topic we investigate the notion of uniform distribution of letters…

Combinatorics · Mathematics 2021-09-01 Hiêp Hàn , Marcos Kiwi , Matías Pavez-Signé

The cornerstone of Boltzmann-Gibbs ($BG$) statistical mechanics is the Boltzmann-Gibbs-Jaynes-Shannon entropy $S_{BG} \equiv -k\int dx f(x)\ln f(x)$, where $k$ is a positive constant and $f(x)$ a probability density function. This theory…

Physics and Society · Physics 2009-11-11 Silvio M. Duarte Queiros , Celia Anteneodo , Constantino Tsallis

We study the variation of word frequencies in Russian literary texts. Our findings indicate that the standard deviation of a word's frequency across texts depends on its average frequency according to a power law with exponent $0.62,$…

Computation and Language · Computer Science 2016-01-20 Vladislav Kargin

We generalize the usual exponential Boltzmann factor to any reasonable and potentially observable distribution function, $B(E)$. By defining generalized logarithms $\Lambda$ as inverses of these distribution functions, we are led to a…

Statistical Mechanics · Physics 2007-05-23 Rudolf Hanel , Stefan Thurner

The analysis of thousands of time series in different languages reveals that word usage presents oscillations with a prevalence of 16-year cycles, mounted on slowly varying trends. These components carry different information: while similar…

Neurons and Cognition · Quantitative Biology 2022-07-20 Alejandro Pardo Pintos , Diego E Shalom , Enzo Tagliazucchi , Gabriel Mindlin , Marcos A Trevisan

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

Probability · Mathematics 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

We model and compute the probability distribution of the letters in random generated words in a language by using the theory of set partitions, Young tableaux and graph theoretical representation methods. This has been of interest for…

Computation and Language · Computer Science 2014-07-24 Alberto Besana , Cristina Martínez

In this paper, a statistical analysis of the structure of one blog community, a kind of social networks, is presented. The quantities such as degree distribution, clustering coefficient, average shortest path length are calculated to…

Statistics Theory · Mathematics 2007-06-13 Feng Fu , Lianghuan Liu , Kai Yang , Long Wang

Despite renewed interest in emergent language simulations with neural networks, little is known about the basic properties of the induced code, and how they compare to human language. One fundamental characteristic of the latter, known as…

Computation and Language · Computer Science 2019-10-16 Rahma Chaabouni , Eugene Kharitonov , Emmanuel Dupoux , Marco Baroni

The rank-size plots of a large number of different physical and socio-economic systems are usually said to follow Zipf's law, but a unique framework for the comprehension of this ubiquitous scaling law is still lacking. Here we show that a…

Physics and Society · Physics 2021-02-03 Giordano De Marzo , Andrea Gabrielli , Andrea Zaccaria , Luciano Pietronero

We consider the number of occurrences of subwords (non-consecutive sub-sequences) in a given word. We first define the notion of subword entropy of a given word that measures the maximal number of occurrences among all possible subwords. We…

Combinatorics · Mathematics 2025-10-06 Wenjie Fang

We develop the information-theoretical concepts required to study the statistical dependencies among three variables. Some of such dependencies are pure triple interactions, in the sense that they cannot be explained in terms of a…

Computation and Language · Computer Science 2015-09-02 Damián G. Hernández , Damián H. Zanette , Inés Samengo

A curious observation was made that the rank statistics of scientific citation numbers follows Zipf-Mandelbrot's law. The same pow-like behavior is exhibited by some simple random citation models. The observed regularity indicates not so…

Physics and Society · Physics 2007-05-23 Z. K. Silagadze

We consider a system composed of a fixed number of particles with total energy smaller or equal to some prescribed value. The particles are non-interacting, indistinguishable and distributed over fixed number of energy levels. The energy…

Probability · Mathematics 2021-03-23 Tomasz M. Łapiński

This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several…

cmp-lg · Computer Science 2008-02-03 Doug Beeferman , Adam Berger , John Lafferty

The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…

Computation and Language · Computer Science 2016-06-23 Christian Bentz , Dimitrios Alikaniotis