Related papers: Big Data and Large Numbers. Interpreting Zipf's La…
Statistical laws describe regular patterns observed in diverse scientific domains, ranging from the magnitude of earthquakes (Gutenberg-Richter law) and metabolic rates in organisms (Kleiber's law), to the frequency distribution of words in…
According to Zipf's meaning-frequency law, words that are more frequent tend to have more meanings. Here it is shown that a linear dependency between the frequency of a form and its number of meanings is found in a family of models of…
The size that an epidemic can reach, measured in terms of the number of fatalities, is an extremely relevant quantity. It has been recently claimed [Cirillo & Taleb, Nature Physics 2020] that the size distribution of major epidemics in…
Deep learning has been successfully applied to various tasks, but its underlying mechanism remains unclear. Neural networks associate similar inputs in the visible layer to the same state of hidden variables in deep layers. The fraction of…
A set of data with positive values follows a Pareto distribution if the log-log plot of value versus rank is approximately a straight line. A Pareto distribution satisfies Zipf's law if the log-log plot has a slope of -1. Since many types…
Present human languages display slightly asymmetric log-normal (Gauss) distribution for size [1-3], whereas present cities follow power law (Pareto-Zipf law)[4]. Our model considers the competition between languages and that between cities…
It is shown that the distribution of low variability periods in the activity of human heart rate typically follows a multi-scaling Zipf's law. The presence or failure of a power law, as well as the values of the scaling exponents, are…
Many scientists are interested in but puzzled by the various inverse power laws with a negative exponent 1 such as the rank-size rule. The rank-size rule is a very simple scaling law followed by many observations of the ubiquitous empirical…
Zipf's law states that the probability of a variable being larger than $s$ is roughly inversely proportional to $s$. In this paper, we evaluate Zipf's law for the distribution of firm size by the number of employees in Brazil. We use…
A family of information theoretic models of communication was introduced more than a decade ago to explain the origins of Zipf's law for word frequencies. The family is a based on a combination of two information theoretic principles:…
As the most fundamental empirical law, Zipf's law has been studied from many aspects. But its meaning is still an open problem. Some models have been constructed to explain Zipf's law. In the letter, a new concept named nonsymmetric entropy…
Power laws arise in a variety of phenomena ranging from matter undergoing phase transition to the distribution of word frequencies in the English language. Usually, their presence is only apparent when data is abundant, and accurately…
Here we sketch a new derivation of Zipf's law for word frequencies based on optimal coding. The structure of the derivation is reminiscent of Mandelbrot's random typing model but it has multiple advantages over random typing: (1) it starts…
The distribution of word probabilities in the monkey model of Zipf's law is associated with two universality properties: (1) the power law exponent converges strongly to $-1$ as the alphabet size increases and the letter probabilities are…
An expression is proposed for determining the error caused on entropy estimates by finite sample effects. This expression is based on the Ansatz that the ranked distribution of probabilities tends to follow an empirical Zipf law.
We consider a system composed of a fixed number of particles with total energy smaller or equal to some prescribed value. The particles are non-interacting, indistinguishable and distributed over fixed number of energy levels. The energy…
The formation of sentences is a highly structured and history-dependent process. The probability of using a specific word in a sentence strongly depends on the 'history' of word-usage earlier in that sentence. We study a simple…
Many complex systems are composed of disparate, interacting types of varying sizes: Species abundances in ecosystems, firm sizes in markets, city populations in countries, word counts in language, etc. A longstanding mystery of complex…
Tokenization is a fundamental step in natural language processing (NLP) and other sequence modeling domains, where the choice of vocabulary size significantly impacts model performance. Despite its importance, selecting an optimal…
This work proves that ranks and shares are statistically dependent on one another, based on simple combinatorics. It presents a formula for rank-share distribution and illustrates that Zipfs law, is descended from expected values of various…