Related papers: Beyond Zipf's Law: The Lavalette Rank Function and…
Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by optimizers such as Adam. These works suggest that the…
Learning from label proportions (LLP) is a generalization of supervised learning in which the training data is available as sets or bags of feature-vectors (instances) along with the average instance-label of each bag. The goal is to train…
Many man-made and natural phenomena, including the intensity of earthquakes, population of cities and size of international wars, are believed to follow power-law distributions. The accurate identification of power-law patterns has…
We show that statistical criticality, i.e. the occurrence of power law frequency distributions, arises in samples that are maximally informative about the underlying generating process. In order to reach this conclusion, we first identify…
Causal processes can give rise to distinctive distributions in the linguistic variables that they affect. Consequently, a secure understanding of a variable's distribution can hold a key to understanding the forces that have causally shaped…
We study the convergence in distribution norms in the Central Limit Theorem for non identical distributed random variables that is $$ \varepsilon_{n}(f):={\mathbb{E}}\Big(f\Big(\frac 1{\sqrt…
Smooth Estimation of probability density and distribution functions from its sample is an attractive and an important problem that has applications in several fields such as, business, medicine, and environment. This article introduces a…
The frequencies at which individual words occur across languages follow power law distributions, a pattern of findings known as Zipf's law. A vast literature argues over whether this serves to optimize the efficiency of human communication,…
We present an overview of possible reasons for the appearance of heavy-tailed distributions in applications to the natural sciences. These distributions include the laws of Pareto, Lotka, and some new ones. The reasons are illustrated using…
Let ${\bf L}$ be the unit exponential random variable and ${\bf Z}_\alpha$ the standard positive $\alpha$-stable random variable. We prove that $\{(1-\alpha) \alpha^{\gamma_\alpha} {\bf Z}_\alpha^{-\gamma_\alpha}, 0< \alpha <1\}$ is…
Logarithmic transformation of the data has been recommended by the literature in the case of highly skewed distributions such as those commonly found in information science. The purpose of the transformation is to make the data conform to…
The Large Deviation Principle (LDP) and the Central Limit Theorem (CLT) are central pillars of probability theory. While their formulations are established under the i.i.d. assumption, the probabilistic foundation for power-law…
Over the last two decades, it has been argued that the Lorentz transformation mechanism, which imposes the generalization of Newton's classical mechanics into Einstein's special relativity, implies a generalization, or deformation, of the…
In his pioneering research, G. K. Zipf observed that more frequent words tend to have more meanings, and showed that the number of meanings of a word grows as the square root of its frequency. He derived this relationship from two…
The present paper is devoted to the relativistic statistical theory, introduced in Phys. Rev. E {\bf 66} (2002) 056125 and Phys. Rev. E {\bf 72} (2005) 036108, predicting the particle distribution function $p(E)= \exp_{\kappa}…
Languages across the world exhibit Zipf's law of abbreviation, namely more frequent words tend to be shorter. The generalized version of the law - an inverse relationship between the frequency of a unit and its magnitude - holds also for…
We propose a new wavelet-based method for density estimation when the data are size-biased. More specifically, we consider a power of the density of interest, where this power exceeds 1/2. Warped wavelet bases are employed, where warping is…
The log-normal distribution is used to describe the positive data, that it has skewed distribution with small mean and large variance. This distribution has application in many sciences for example medicine, economics, biology and…
Zipf's law is shown to arise as the variational solution of a problem formulated in Fisher's terms. An appropriate minimization process involving Fisher information and scale-invariance yields this universal rank distribution. As an example…
Zipf's law defines an inverse proportion between a word's ranking in a given corpus and its frequency in it, roughly dividing the vocabulary into frequent words and infrequent ones. Here, we stipulate that within a domain an author's…