Related papers: Using the Gini coefficient to characterize the sha…
Over the past decades, researchers and ML practitioners have come up with better and better ways to build, understand and improve the quality of ML models, but mostly under the key assumption that the training data is distributed…
Distributed learning facilitates the scaling-up of data processing by distributing the computational burden over several nodes. Despite the vast interest in distributed learning, generalization performance of such approaches is not well…
This work explores the distribution of citations for the publications of top scientists. A first objective is to find out whether the 80-20 Pareto rule applies, that is if 80% of the citations to a top scientist's work concern 20% of their…
Every practical method to solve the Schr\"odinger equation for interacting many-particle systems introduces approximations. Such methods are therefore plagued by systematic errors. For computational chemistry, it is decisive to quantify the…
We examine the efficiency of the mean deviation and Gini's mean difference (the mean of all pairwise distances). Our findings support the viewpoint that Gini's mean difference combines the advantages of the mean deviation and the standard…
Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons…
Many engineering systems are subject to spatially distributed uncertainty, i.e. uncertainty that can be modeled as a random field. Altering the mean or covariance of this uncertainty will in general change the statistical distribution of…
Statistical inference on the explained variation of an outcome by a set of covariates is of particular interest in practice. When the covariates are of moderate to high-dimension and the effects are not sparse, several approaches have been…
The categorical Gini correlation is an alternative measure of dependence between a categorical and numerical variables, which characterizes the independence of the variables. A nonparametric test for the equality of K distributions has been…
Given a random variable $X$ and considered a family of its possible distortions, we define two new measures of distance between $X$ and each its distortion. For these distance measures, which are extensions of the Gini's mean difference,…
This review is designed to introduce mathematicians and computational scientists to quantum computing (QC) through the lens of uncertainty quantification (UQ) by presenting a mathematically rigorous and accessible narrative for…
We consider a two-component mixture model with one known component. We develop methods for estimating the mixing proportion and the unknown distribution nonparametrically, given i.i.d.~data from the mixture model, using ideas from shape…
Missing data imputation (MDI) is a fundamental problem in many scientific disciplines. Popular methods for MDI use global statistics computed from the entire data set (e.g., the feature-wise medians), or build predictive models operating…
In this paper, we discuss the worst-case of distortion riskmetrics for general distributions when only partial information (mean and variance) is known. This result is applicable to general class of distortion risk measures and variability…
Chemical toxicity prediction using machine learning is important in drug development to reduce repeated animal and human testing, thus saving cost and time. It is highly recommended that the predictions of computational toxicology models…
The Gini index underestimates inequality for heavy-tailed distributions: for example, a Pareto distribution with exponent 1.5 (which has infinite variance) has the same Gini index as any exponential distribution (a mere 0.5). This is…
In this paper we will show that the Gini coefficient and the introduced measure of angular inequality are special cases of a wider indexed family of measurements. We will discuss the properties of the defined class based, inter alia, on a…
It is not unusual for a data analyst to encounter data sets distributed across several computers. This can happen for reasons such as privacy concerns, efficiency of likelihood evaluations, or just the sheer size of the whole data set. This…
Linear mixed-effects models are widely used in analyzing clustered or repeated measures data. We propose a quasi-likelihood approach for estimation and inference of the unknown parameters in linear mixed-effects models with high-dimensional…
Computational molecular modeling and visualization has seen significant progress in recent years with sev- eral molecular modeling and visualization software systems in use today. Nevertheless the molecular biology community lacks…