Related papers: Rank Distributions: Frequency vs. Magnitude
Understanding the innovation process, that is the underlying mechanisms through which novelties emerge, diffuse and trigger further novelties is undoubtedly of fundamental importance in many areas (biology, linguistics, social science and…
The prevailing maximum likelihood estimators for inferring power law models from rank-frequency data are biased. The source of this bias is an inappropriate likelihood function. The correct likelihood function is derived and shown to be…
In solar physics, especially in exploratory stages of research, it is often necessary to compare the power spectra of two or more time series. One may, for instance, wish to estimate what the power spectrum of the combined data sets might…
Power-law distributions with various exponents are studied. We first introduce a simple and generic model that reproduces Zipf's law. We can regard this model both as the time evolution of the population of cities and that of the asset…
A wide range of natural and engineered phenomena rely on large networks of interacting units to reach a dynamical consensus state where the system collectively operates. Here we study the dynamics of self-organizing systems and show that…
This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test…
Stochastic equations constitute a major ingredient in many branches of science, from physics to biology and engineering. Not surprisingly, they appear in many quantitative studies of complex systems. In particular, this type of equation is…
In a language corpus, the probability that a word occurs $n$ times is often proportional to $1/n^2$. Assigning rank, $s$, to words according to their abundance, $\log s$ vs $\log n$ typically has a slope of minus one. That simple Zipf's law…
We investigate the degree distribution resulting from graph generation models based on rank-based attachment. In rank-based attachment, all vertices are ranked according to a ranking scheme. The link probability of a given vertex is…
The general relationship between an arbitrary frequency distribution and the expectation value of the frequency distributions of its samples is esablished. A set of combinations of expectation values whose value does not in general depend…
Stopwords are words that are not very informative to the content or the meaning of a language text. Most stopwords are function words but can also be common verbs, adjectives and adverbs. In contrast to the well known Zipf's law for…
The distribution of inter-occurrence time between seismic events is a quantity of great interest in seismic risk assessment. We evaluate this distribution for different models of earthquakes occurrence and follow two distinct approaches:…
We show how generalized Gibbs-Shannon entropies can provide new insights on the statistical properties of texts. The universal distribution of word frequencies (Zipf's law) implies that the generalized entropies, computed at the word level,…
We consider a new functional inequality controlling the rate of relative entropy decay for random walks, the interchange process and more general block-type dynamics for permutations. The inequality lies between the classical logarithmic…
In this paper, we study two large data sets containing the information of two different human behaviors: blog-posting and wiki-revising. In both cases, the interevent time distributions decay as power-laws at both individual and population…
The `plate tectonics' is an observed fact and most models of earthquake incorporate that through the frictional dynamics (stick-slip) of two surfaces where one surface moves over the other. These models are more or less successful to…
Phoneme frequency distributions exhibit robust statistical regularities across languages, including exponential-tailed rank-frequency patterns and a negative relationship between phonemic inventory size and the relative entropy of the…
Zipf's power-law distribution is a generic empirical statistical regularity found in many complex systems. However, rather than universality with a single power-law exponent (equal to 1 for Zipf's law), there are many reported deviations…
The distribution of numbers in human documents is determined by a variety of diverse natural and human factors, whose relative significance can be evaluated by studying the numbers' frequency of occurrence. Although it has been studied…
English words and the outputs of many other natural processes are well-known to follow a Zipf distribution. Yet this thoroughly-established property has never been shown to help compress or predict these important processes. We show that…