Related papers: Power Law in a Bounded Range: Estimating the Lower…
It is well known that the entropy $H(X)$ of a discrete random variable $X$ is always greater than or equal to the entropy $H(f(X))$ of a function $f$ of $X$, with equality if and only if $f$ is one-to-one. In this paper, we give tight…
Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggests that reweighting or curating data towards a uniform distribution may help models…
Let $(d_n)$ be a sequence of positive numbers and let $(X_n)$ be a sequence of positive independent random variables. We provide an upper bound for the deviation between the distribution of the mantissaes of $(X_n^{d_n})$ and the Benford's…
Recent theories suggest that Neural Scaling Laws arise whenever the task is linearly decomposed into power-law distributed units. Alternatively, scaling laws also emerge when data exhibit a hierarchically compositional structure, as is…
Exact upper and lower bounds on the ratio $\mathsf{E}w(\mathbf{X}-\mathbf{v})/\mathsf{E}w(\mathbf{X})$ for a centered Gaussian random vector $\mathbf{X}$ in $\mathbb{R}^n$, as well as bounds on the rate of change of…
Three versions of the Weak Law of Large Numbers are proposed for weakly dependent and generally speaking non-equally distributed random variables, with finite or possibly infinite expectations.
Finite-volume extrapolation is an important step for extracting physical observables from lattice calculations. However, it is a significant challenge for the system with long-range interactions. We employ symbolic regression to regress…
In this article we derive the best possible upper bound for $E[\max{X_i}-\min_i{X_i}]$ under given means and variances on $n$ random variables $X_i$. The random vector $(X_1,...,X_n)$ is allowed to have any dependence structure, provided $E…
In applied probability, the normal approximation is often used for the distribution of data with assumed additive structure. This tradition is based on the central limit theorem for sums of (independent) random variables. However, it is…
The power law is useful in describing count phenomena such as network degrees and word frequencies. With a single parameter, it captures the main feature that the frequencies are linear on the log-log scale. Nevertheless, there have been…
The size or energy of diverse structures or phenomena in geoscience appears to follow power-law distributions. A rigorous statistical analysis of such observations is tricky, though. Observables can span several orders of magnitude, but the…
We study the emergence of a power law distribution in the systems which can be characterized by a hierarchically organized supplying network. It is shown that conservation laws on the branches of the network can, at some approximation,…
We derive new upper and lower bounds for probabilities that $r$ or at least $r$ from $n$ events occur. These bounds can turn to equalities. The method is discussed as well. It works for measurable space and measures with sign, too. We also…
Let I_1,...,I_n be independent but not necessarily identically distributed Bernoulli random variables, and let X_n=\sum_{j=1}^nI_j. For \nu in a bounded region, a local central limit theorem expansion of P(X_n=EX_n+\nu) is developed to any…
Power law-like size distributions are ubiquitous in astrophysical instabilities. There are at least four natural effects that cause deviations from ideal power law size distributions, which we model here in a generalized way: (1) a physical…
When the probability of measuring a particular value of some quantity varies inversely as a power of that value, the quantity is said to follow a power law, also known variously as Zipf's law or the Pareto distribution. Power laws appear…
This paper defines theoretical lower bounds of uncertainty of observations of macroeconomic variables that depend on statistical moments and correlations of random values and volumes of market trades. Any econometric assessments of…
I present an analytic method for estimating the errors in fitting a distribution. A well-known theorem from statistics gives the minimum variance bound (MVB) for the uncertainty in estimating a set of parameters $\l_i$, when a distribution…
The traditional lower bound estimation method for powerlaw distributions based on the Kolmogorov-Smirnov distance proved to perform better than other competing methods. However, if applied to very large collections of data, such a method…
Scale-free networks play a fundamental role in the study of complex networks and various applied fields due to their ability to model a wide range of real-world systems. A key characteristic of these networks is their degree distribution,…