Related papers: How to (Not) Estimate Gini Coefficients for Fat Ta…
Social inequality manifested across different strata of human existence can be quantified in several ways. Here we compute non-entropic measures of inequality such as Lorenz curve, Gini index and the recently introduced $k$ index…
We consider the probability that a weighted sum of $n$ i.i.d. random variables $X_j$, $j = 1, . . ., n$, with stretched exponential tails is larger than its expectation and determine the rate of its decay, under suitable conditions on the…
Measuring distances in a multidimensional setting is a challenging problem, which appears in many fields of science and engineering. In this paper, to measure the distance between two multivariate distributions, we introduce a new measure…
How reliably can we trust the scores obtained from social bias benchmarks as faithful indicators of problematic social biases in a given language model? In this work, we study this question by contrasting social biases with non-social…
In this paper, we discuss the worst-case of distortion riskmetrics for general distributions when only partial information (mean and variance) is known. This result is applicable to general class of distortion risk measures and variability…
Machine learning models trained on real-world data may inadvertently make biased predictions that negatively impact marginalized communities. Reweighting, which assigns a weight to each data point used during model training, can mitigate…
We study large deviation properties of probability distributions with either a compact support or a fat tail by comparing them with q-deformed exponential distributions. Our main result is a large deviation property for probability…
A novel statistical method is proposed and investigated for estimating a heavy tailed density under mild smoothness assumptions. Statistical analyses of heavy-tailed distributions are susceptible to the problem of sparse information in the…
In this article two methods to distinguish between polynomial and exponential tails are introduced. The methods are mainly based on the properties of the residual coefficient of variation for the exponential and non-exponential…
A wide range of natural and social phenomena result in observables whose distributions can be well approximated by a power-law decay. The well-known Hill estimator of the tail exponent provides results which are in many respects superior to…
The most popular approach in extreme value statistics is the modelling of threshold exceedances using the asymptotically motivated generalised Pareto distribution. This approach involves the selection of a high threshold above which the…
This paper contributes to answering a question that is of crucial importance in risk management and extreme value theory: How to select the threshold above which one assumes that the tail of a distribution follows a generalized Pareto…
Length-based methods are the cornerstone of many population studies and stock assessments. This study tested two widely used methods: the Powell-Wetherall (P-W) plot and the Lmax approach, i.e., estimating Linf directly from Lmax. In most…
While fat-tailed densities commonly arise as posterior and marginal distributions in robust models and scale mixtures, they present challenges when Gaussian-based variational inference fails to capture tail decay accurately. We first…
In risk management, tail risks are of crucial importance. The assessment of risks should be carried out in accordance with the regulatory authority's requirement at high quantiles. In general, the underlying distribution function is…
Heavy-tailed metrics are common and often critical to product evaluation in the online world. While we may have samples large enough for Central Limit Theorem to kick in, experimentation is challenging due to the wide confidence interval of…
We study probability inequalities leading to tail estimates in a general semigroup $\mathscr{G}$ with a translation-invariant metric $d_{\mathscr{G}}$. (An important and central example of this in the functional analysis literature is that…
The simple linear model $$Y_i = \alpha + \beta \, x_i + \epsilon_i \qquad i=1,2, \ldots,N \geq 2$$ is considered, where the $x_i$'s are given constants and $\epsilon_1, \epsilon_2 , \ldots, \epsilon_N$ are iid with continuous distribution…
Given a random variable $X$ and considered a family of its possible distortions, we define two new measures of distance between $X$ and each its distortion. For these distance measures, which are extensions of the Gini's mean difference,…
Gaussian noise injections (GNIs) are a family of simple and widely-used regularisation methods for training neural networks, where one injects additive or multiplicative Gaussian noise to the network activations at every iteration of the…