English
Related papers

Related papers: Using the Gini coefficient to characterize the sha…

200 papers

When machine learning models are deployed on a test distribution different from the training distribution, they can perform poorly, but overestimate their performance. In this work, we aim to better estimate a model's performance under…

Machine Learning · Computer Science 2020-07-08 Ching-Yao Chuang , Antonio Torralba , Stefanie Jegelka

Despite the popularity of feature importance (FI) measures in interpretable machine learning, the statistical adequacy of these methods is rarely discussed. From a statistical perspective, a major distinction is between analyzing a…

Machine Learning · Statistics 2023-05-03 Kristin Blesch , David S. Watson , Marvin N. Wright

We propose two bounded comparison metrics that may be implemented to arbitrary dimensions in regression tasks. One quantifies the structure of uncertainty and the other quantifies the distribution of uncertainty. The structure metric…

Machine Learning · Computer Science 2022-03-10 Ethan Pickering , Themistoklis P. Sapsis

A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label…

Machine Learning · Statistics 2021-11-17 Nilesh Tripuraneni , Ben Adlam , Jeffrey Pennington

Numerical models based on partial differential equations (PDE), or integro-differential equations, are ubiquitous in engineering and science, making it possible to understand or design systems for which physical experiments would be…

Computational Physics · Physics 2021-04-02 Julien Bect , Souleymane Zio , Guillaume Perrin , Claire Cannamela , Emmanuel Vazquez

Alpha-based performance evaluation may fail to capture correlated residuals due to model errors. This paper proposes using the Generalized Information Ratio (GIR) to measure performance under misspecified benchmarks. Motivated by the…

Portfolio Management · Quantitative Finance 2018-04-24 Zhongzhi Lawrence He

The mixture of Gaussian distributions, a soft version of k-means , is considered a state-of-the-art clustering algorithm. It is widely used in computer vision for selecting classes, e.g., color, texture, and shapes. In this algorithm, each…

Machine Learning · Statistics 2016-12-30 Mahajabin Rahman , Davi Geiger

Uncertainty estimation, which provides a means of building explainable neural networks for medical imaging applications, have mostly been studied for single deep learning models that focus on a specific task. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Leonhard F. Feiner , Martin J. Menten , Kerstin Hammernik , Paul Hager , Wenqi Huang , Daniel Rueckert , Rickmer F. Braren , Georgios Kaissis

We describe a method for fitting distributions to data which only requires knowledge of the parametric form of either the signal or the background but not both. The unknown distribution is fit using a non-parametric kernel density…

Data Analysis, Statistics and Probability · Physics 2015-06-03 Wolfgang A. Rolke , Angel M. López

Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment planning. These biomarkers are typically estimated from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiameng Li , Teodora Popordanoska , Aleksei Tiulpin , Sebastian G. Gruber , Frederik Maes , Matthew B. Blaschko

If a discrete probability distribution in a model being tested for goodness-of-fit is not close to uniform, then forming the Pearson chi-square statistic can involve division by nearly zero. This often leads to serious trouble in practice…

Methodology · Statistics 2011-09-16 William Perkins , Mark Tygert , Rachel Ward

The brilliant method due to Good and Turing allows for estimating objects not occurring in a sample. The problem, known under names "sample coverage" or "missing mass" goes back to their cryptographic work during WWII, but over years has…

Machine Learning · Statistics 2021-04-16 Maciej Skorski

Semiparametric models are useful in econometrics, social sciences and medicine application. In this paper, a new estimator based on least square methods is proposed to estimate the direction of unknown parameters in semi-parametric models.…

Methodology · Statistics 2023-03-10 Jinyue Han , Jun Wang , Wei Gao , Man-Lai Tang

Diffusion models have recently emerged as powerful learners for simulation-based inference (SBI), enabling fast and accurate estimation of latent parameters from simulated and real data. Their score-based formulation offers a flexible way…

Machine Learning · Statistics 2026-01-30 Jonas Arruda , Niels Bracher , Ullrich Köthe , Jan Hasenauer , Stefan T. Radev

Statistical inference from high-dimensional data with low-dimensional structures has recently attracted lots of attention. In machine learning, deep generative modeling approaches implicitly estimate distributions of complex objects by…

Statistics Theory · Mathematics 2022-02-21 Rong Tang , Yun Yang

In this paper, we introduce a novel flexible Gini index, referred to as the extended Gini index, which is defined through ordered differences between the $j$th and $k$th order statistics within subsamples of size $m$, for indices satisfying…

Methodology · Statistics 2025-05-06 Roberto Vila , Helton Saulo

We propose a new family of inequality indices that bridges the Hoover index and the Gini coefficient. The measure is defined as the normalized expected absolute value of a convex combination of deviations from the mean and pairwise…

Methodology · Statistics 2026-03-30 Roberto Vila , Helton Saulo , Felipe Quintino

Major drawback of studying diffusion in multi-component systems is the lack of suitable techniques to estimate the diffusion parameters. In this study, a generalized treatment to determine the intrinsic diffusion coefficients in…

Materials Science · Physics 2022-08-23 Sangeeta Santra , Aloke Paul

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in non-standard settings, particularly those where observations are not independently and…

The Gini index signals only the dispersion of the distribution and is not very sensitive to income differences at the tails of the distribution. The widely used index of inequality can be adjusted to also measure distributional asymmetry by…

Econometrics · Economics 2022-09-15 Mario Schlemmer