Related papers: Finding trends and statistical patterns in name me…
More than one billion data sampled with different frequencies from several financial instruments were investigated with the aim of testing whether they involve power law. As a result, a known power law with the power exponent around -4 was…
We analyze time evolution of statistical distributions of citations to scientific papers published in one year. While these distributions can be fitted by a power-law dependence we find that they are nonstationary and the exponent of the…
Scientific cooperation on an international level has been well studied in the literature. However, much less is known about this cooperation on the intercontinental level. In this paper, we address this issue by creating a collection of…
This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual word number on text…
The frequency distributions of DNA k-mers are shaped by fundamental biological processes and offer a window into genome structure and evolution. Inspired by analogies to natural language, prior studies have attempted to model genomic k-mer…
Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…
This paper explores the relationship between the inner economical structure of communities and their population distribution through a rank-rank analysis of official data, along statistical physics ideas within two techniques. The data is…
Users of social networking services construct their personal social networks by creating asymmetric and symmetric social links. Users usually follow friends and selected famous entities that include celebrities and news agencies. In this…
The paper concerns the rates of power-law growth of mutual information computed for a stationary measure or for a universal code. The rates are called Hilberg exponents and four such quantities are defined for each measure and each code:…
We investigate transitions of portals users between different subpages. A weighted network of portals subpages is reconstructed where edge weights are numbers of corresponding transitions. Distributions of link weights and node strengths…
Statistical regularities in human language have fascinated researchers for decades, suggesting deep underlying principles governing its evolution and information structuring for efficient communication. While Zipf's Law describes the…
Zipf's law states that the frequency of an observation with a given value is inversely proportional to the square of that value; Taylor's law, instead, describes the scaling between fluctuations in the size of a population and its mean.…
When the probability of measuring a particular value of some quantity varies inversely as a power of that value, the quantity is said to follow a power law, also known variously as Zipf's law or the Pareto distribution. Power laws appear…
In this second part of our survey on the social and natural distributions, we investigate some models, which intend to explain the statistical regularity of the natural and social distributions. There is a large variety of models and in…
Throughout history most young adults have chosen to live where their parents did while a smaller number moved away. This is sufficient, by proof and simulation, to account for the well-known power law distributions of city sizes. The model…
Present human languages display slightly asymmetric log-normal (Gauss) distribution for size [1-3], whereas present cities follow power law (Pareto-Zipf law)[4]. Our model considers the competition between languages and that between cities…
It turns out that some empirical facts in Big Data are the effects of properties of large numbers. Zipf's law 'noise' is an example of such an artefact. We expose several properties of the power law distributions and of similar distribution…
Large language models with a huge number of parameters, when trained on near internet-sized number of tokens, have been empirically shown to obey neural scaling laws: specifically, their performance behaves predictably as a power law in…
We have translated fractional Brownian motion (FBM) signals into a text based on two ''letters'', as if the signal fluctuations correspond to a constant stepsize random walk. We have applied the Zipf method to extract the $\zeta '$ exponent…
Elections, specially in countries such as Brazil with an electorate of the order of 100 million people, yield large-scale data-sets embodying valuable information on the dynamics through which individuals influence each other and make…