Related papers: Gini estimation under infinite variance
The tail index, indicating the degree of fatness of the tail distribution, is an important component of extreme value theory since it dominates the asymptotic distribution of extreme values such as the sample maximum. In this paper, we…
The simple linear model $$Y_i = \alpha + \beta \, x_i + \epsilon_i \qquad i=1,2, \ldots,N \geq 2$$ is considered, where the $x_i$'s are given constants and $\epsilon_1, \epsilon_2 , \ldots, \epsilon_N$ are iid with continuous distribution…
The pursuit of having an appropriate level of income inequality should be viewed as one of the biggest challenges facing academic scholars as well as policy makers. Unfortunately, research on this issue is currently lacking. This study is…
The Gini index signals only the dispersion of the distribution and is not very sensitive to income differences at the tails of the distribution. The widely used index of inequality can be adjusted to also measure distributional asymmetry by…
The heavy reliance on data is one of the major reasons that currently limit the development of deep learning. Data quality directly dominates the effect of deep learning models, and the long-tailed distribution is one of the factors…
The subject of robust estimation in time series is widely discussed in literature. One of the approaches is to use GM-estimation. This method incorporates a broad class of nonparametric estimators which under suitable conditions includes…
Grouped data in form of income shares have been conventionally used to estimate income inequality due to the lack of availability of individual records. Most prior research on economic inequality relies on lower bounds of inequality…
The study of loss function distributions is critical to characterize a model's behaviour on a given machine learning problem. For example, while the quality of a model is commonly determined by the average loss assessed on a testing set,…
In this paper we are concerned with the analysis of heavy-tailed data when a portion of the extreme values is unavailable. This research was motivated by an analysis of the degree distributions in a large social network. The degree…
We suggest a simple Gaussian mixture model for data generation that complies with Feldman's long tail theory (2020). We demonstrate that a linear classifier cannot decrease the generalization error below a certain level in the proposed…
The entropic risk measure is widely used in high-stakes decision-making across economics, management science, finance, and safety-critical control systems because it captures tail risks associated with uncertain losses. However, when data…
Estimating extreme quantiles is an important task in many applications, including financial risk management and climatology. More important than estimating the quantile itself is to insure zero coverage error, which implies the quantile…
Estimation of genewise variance arises from two important applications in microarray data analysis: selecting significantly differentially expressed genes and validation tests for normalization of microarray data. We approach the problem by…
Generative Adversarial Networks (GAN) are a powerful methodology and can be used for unsupervised anomaly detection, where current techniques have limitations such as the accurate detection of anomalies near the tail of a distribution. GANs…
We present the elliptical processes -- a family of non-parametric probabilistic models that subsumes the Gaussian process and the Student-t process. This generalization includes a range of new fat-tailed behaviors yet retains computational…
We extend biologically-informed neural networks (BINNs) for genomic prediction (GP) and selection (GS) in crops by integrating thousands of single-nucleotide polymorphisms (SNPs) with multi-omics measurements and prior biological knowledge.…
The goal of this presentation is to build an efficient non-parametric Bayes classifier in the presence of large numbers of predictors. When analyzing such data, parametric models are often too inflexible while non-parametric procedures tend…
We prove tail estimates for variables $\sum_i f(X_i)$, where $(X_i)_i$ is the trajectory of a random walk on an undirected graph (or, equivalently, a reversible Markov chain). The estimates are in terms of the maximum of the function $f$,…
We analyze neural scaling laws in a solvable model of last-layer fine-tuning where targets have intrinsic, instance-heterogeneous difficulty. In our Latent Instance Difficulty (LID) model, each input's target variance is governed by a…
We consider the problem of inference for non-stationary time series with heavy-tailed error distribution. Under a time-varying linear process framework we show that there exists a suitable local approximation by a stationary process with…