English
Related papers

Related papers: Robust analogs to the Coefficient of Variation

200 papers

Standard practice obtains an unbiased variance estimator by dividing by $N-1$ rather than $N$. Yet if only half the data are used to compute the mean, dividing by $N$ can still yield an unbiased estimator. We show that an alternative mean…

Statistics Theory · Mathematics 2025-04-10 Dai Akita

Many modern data analyses benefit from explicitly modeling dependence structure in data -- such as measurements across time or space, ordered words in a sentence, or genes in a genome. A gold standard evaluation technique is structured…

Machine Learning · Statistics 2020-12-02 Soumya Ghosh , William T. Stephenson , Tin D. Nguyen , Sameer K. Deshpande , Tamara Broderick

The most common way to sample from a probability distribution is to use Monte-Carlo methods. For distributions on a continuous state space, one can find diffusions with the target distribution as equilibrium measure, so that the state of…

Probability · Mathematics 2015-10-28 Chii-Ruey Hwang , Raoul Normand , Sheng-Jhih Wu

Covariance estimation becomes challenging in the regime where the number p of variables outstrips the number n of samples available to construct the estimate. One way to circumvent this problem is to assume that the covariance matrix is…

Probability · Mathematics 2012-06-14 Richard Y. Chen , Alex Gittens , Joel A. Tropp

This short note proposes two additive corrections to a pair of relations published in Wan et al. in order to extend them to a small sample size condition. In particular we focus the interest on the possibility to provide an estimate to the…

Applications · Statistics 2023-10-20 Massimo Borelli

The maximum mean discrepancy (MMD) is a kernel-based distance between probability distributions useful in many applications (Gretton et al. 2012), bearing a simple estimator with pleasing computational and statistical properties. Being able…

Machine Learning · Statistics 2022-11-16 Danica J. Sutherland , Namrata Deka

A distributed estimation scheme where the sensors transmit with constant modulus signals over a multiple access channel is considered. The proposed estimator is shown to be strongly consistent for any sensing noise distribution in the…

Information Theory · Computer Science 2015-05-14 Cihan Tepedelenlioglu , Adarsh B. Narasimhamurthy

Estimating a distribution given access to its unnormalized density is pivotal in Bayesian inference, where the posterior is generally known only up to an unknown normalizing constant. Variational inference and Markov chain Monte Carlo…

Machine Learning · Statistics 2025-05-06 Daniel Ward , Mark Beaumont , Matteo Fasiolo

The distance covariance of two random vectors is a measure of their dependence. The empirical distance covariance and correlation can be used as statistical tools for testing whether two random vectors are independent. We propose an analogs…

Statistics Theory · Mathematics 2017-03-31 Muneya Matsui , Thomas Mikosch , Gennady Samorodnitsky

This paper is devoted to the estimators of the mean that provide strong non-asymptotic guarantees under minimal assumptions on the underlying distribution. The main ideas behind proposed techniques are based on bridging the notions of…

Statistics Theory · Mathematics 2019-05-07 Stanislav Minsker

(To appear in The American Statistician.) Distance covariance (Sz\'ekely, Rizzo, and Bakirov, 2007) is a fascinating recent notion, which is popular as a test for dependence of any type between random variables $X$ and $Y$. This approach…

Methodology · Statistics 2024-07-08 Jakob Raymaekers , Peter J. Rousseeuw

The conditional value-at-risk (CVaR) is a useful risk measure in fields such as machine learning, finance, insurance, energy, etc. When measuring very extreme risk, the commonly used CVaR estimation method of sample averaging does not work…

Methodology · Statistics 2021-03-10 Dylan Troop , Frédéric Godin , Jia Yuan Yu

The term moderate deviations is often used in the literature to mean a class of large deviation principles that, in some sense, fills the gap between a convergence in probability of some random variables to a constant and a weak convergence…

Probability · Mathematics 2024-11-20 Rita Giuliano , Claudio Macci , Barbara Pacchiarotti

A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label…

Machine Learning · Statistics 2021-11-17 Nilesh Tripuraneni , Ben Adlam , Jeffrey Pennington

Several new geometric quantile-based measures for multivariate dispersion, skewness, kurtosis, and spherical asymmetry are defined. These measures differ from existing measures, which use volumes and are easy to calculate. Some theoretical…

Statistics Theory · Mathematics 2024-12-30 Ha-Young Shin , Hee-Seok Oh

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so…

Computation and Language · Computer Science 2018-06-20 Henry B. Moss , David S. Leslie , Paul Rayson

Cross-validation (CV) is a common method to tune machine learning methods and can be used for model selection in regression as well. Because of the structured nature of small, traditional experimental designs, the literature has warned…

Applications · Statistics 2025-06-18 Maria L. Weese , Byran J. Smucker , David J. Edwards

We consider the problem of estimating the mean of a random vector based on i.i.d. observations and adversarial contamination. We introduce a multivariate extension of the trimmed-mean estimator and show its optimal performance under minimal…

Statistics Theory · Mathematics 2020-02-25 Gabor Lugosi , Shahar Mendelson

Cross-validation (CV) is a technique for evaluating the ability of statistical models/learning systems based on a given data set. Despite its wide applicability, the rather heavy computational cost can prevent its use as the system size…

Machine Learning · Statistics 2016-10-26 Yoshiyuki Kabashima , Tomoyuki Obuchi , Makoto Uemura

In Bayesian analysis, the posterior follows from the data and a choice of a prior and a likelihood. One hopes that the posterior is robust to reasonable variation in the choice of prior and likelihood, since this choice is made by the…

Methodology · Statistics 2015-12-09 Ryan Giordano , Tamara Broderick , Michael Jordan
‹ Prev 1 4 5 6 7 8 10 Next ›