Related papers: From Boltzmann to Zipf through Shannon and Jaynes
Biosemiosis is a process of choice-making between simultaneously alternative options. It is well-known that, when sufficiently young children encounter a new word, they tend to interpret it as pointing to a meaning that does not have a word…
The Zipf distribution is a probability distribution widely used by scientists from various disciplines due to its ubiquity. Some of these areas include linguistics, physics, genetics, and sociology, among others. In this paper, it is proved…
In the Yule-Simon process, selection of words follows the preferential attachment mechanism, resulting in the power-law growth in the cumulative number of individual word occurrences. This is derived using mean-field approximation, assuming…
Zipf's law implies the statistical distributions of hyperbolic type, which can describe the properties of stability and entropy loss in linguistics. We present the information theory from which follows that if the system is described by…
Words are sequences of letters over a finite alphabet. We study two intimately related topics for this object: quasi-randomness and limit theory. With respect to the first topic we investigate the notion of uniform distribution of letters…
The cornerstone of Boltzmann-Gibbs ($BG$) statistical mechanics is the Boltzmann-Gibbs-Jaynes-Shannon entropy $S_{BG} \equiv -k\int dx f(x)\ln f(x)$, where $k$ is a positive constant and $f(x)$ a probability density function. This theory…
We study the variation of word frequencies in Russian literary texts. Our findings indicate that the standard deviation of a word's frequency across texts depends on its average frequency according to a power law with exponent $0.62,$…
We generalize the usual exponential Boltzmann factor to any reasonable and potentially observable distribution function, $B(E)$. By defining generalized logarithms $\Lambda$ as inverses of these distribution functions, we are led to a…
The analysis of thousands of time series in different languages reveals that word usage presents oscillations with a prevalence of 16-year cycles, mounted on slowly varying trends. These components carry different information: while similar…
Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…
We model and compute the probability distribution of the letters in random generated words in a language by using the theory of set partitions, Young tableaux and graph theoretical representation methods. This has been of interest for…
In this paper, a statistical analysis of the structure of one blog community, a kind of social networks, is presented. The quantities such as degree distribution, clustering coefficient, average shortest path length are calculated to…
Despite renewed interest in emergent language simulations with neural networks, little is known about the basic properties of the induced code, and how they compare to human language. One fundamental characteristic of the latter, known as…
The rank-size plots of a large number of different physical and socio-economic systems are usually said to follow Zipf's law, but a unique framework for the comprehension of this ubiquitous scaling law is still lacking. Here we show that a…
We consider the number of occurrences of subwords (non-consecutive sub-sequences) in a given word. We first define the notion of subword entropy of a given word that measures the maximal number of occurrences among all possible subwords. We…
We develop the information-theoretical concepts required to study the statistical dependencies among three variables. Some of such dependencies are pure triple interactions, in the sense that they cannot be explained in terms of a…
A curious observation was made that the rank statistics of scientific citation numbers follows Zipf-Mandelbrot's law. The same pow-like behavior is exhibited by some simple random citation models. The observed regularity indicates not so…
We consider a system composed of a fixed number of particles with total energy smaller or equal to some prescribed value. The particles are non-interacting, indistinguishable and distributed over fixed number of energy levels. The energy…
This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several…
The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…