中文
相关论文

相关论文: Robust analogs to the Coefficient of Variation

200 篇论文

Standard practice obtains an unbiased variance estimator by dividing by $N-1$ rather than $N$. Yet if only half the data are used to compute the mean, dividing by $N$ can still yield an unbiased estimator. We show that an alternative mean…

统计理论 · 数学 2025-04-10 Dai Akita

Many modern data analyses benefit from explicitly modeling dependence structure in data -- such as measurements across time or space, ordered words in a sentence, or genes in a genome. A gold standard evaluation technique is structured…

The most common way to sample from a probability distribution is to use Monte-Carlo methods. For distributions on a continuous state space, one can find diffusions with the target distribution as equilibrium measure, so that the state of…

概率论 · 数学 2015-10-28 Chii-Ruey Hwang , Raoul Normand , Sheng-Jhih Wu

Covariance estimation becomes challenging in the regime where the number p of variables outstrips the number n of samples available to construct the estimate. One way to circumvent this problem is to assume that the covariance matrix is…

概率论 · 数学 2012-06-14 Richard Y. Chen , Alex Gittens , Joel A. Tropp

This short note proposes two additive corrections to a pair of relations published in Wan et al. in order to extend them to a small sample size condition. In particular we focus the interest on the possibility to provide an estimate to the…

应用统计 · 统计学 2023-10-20 Massimo Borelli

The maximum mean discrepancy (MMD) is a kernel-based distance between probability distributions useful in many applications (Gretton et al. 2012), bearing a simple estimator with pleasing computational and statistical properties. Being able…

机器学习 · 统计学 2022-11-16 Danica J. Sutherland , Namrata Deka

A distributed estimation scheme where the sensors transmit with constant modulus signals over a multiple access channel is considered. The proposed estimator is shown to be strongly consistent for any sensing noise distribution in the…

信息论 · 计算机科学 2015-05-14 Cihan Tepedelenlioglu , Adarsh B. Narasimhamurthy

Estimating a distribution given access to its unnormalized density is pivotal in Bayesian inference, where the posterior is generally known only up to an unknown normalizing constant. Variational inference and Markov chain Monte Carlo…

机器学习 · 统计学 2025-05-06 Daniel Ward , Mark Beaumont , Matteo Fasiolo

The distance covariance of two random vectors is a measure of their dependence. The empirical distance covariance and correlation can be used as statistical tools for testing whether two random vectors are independent. We propose an analogs…

统计理论 · 数学 2017-03-31 Muneya Matsui , Thomas Mikosch , Gennady Samorodnitsky

This paper is devoted to the estimators of the mean that provide strong non-asymptotic guarantees under minimal assumptions on the underlying distribution. The main ideas behind proposed techniques are based on bridging the notions of…

统计理论 · 数学 2019-05-07 Stanislav Minsker

(To appear in The American Statistician.) Distance covariance (Sz\'ekely, Rizzo, and Bakirov, 2007) is a fascinating recent notion, which is popular as a test for dependence of any type between random variables $X$ and $Y$. This approach…

统计方法学 · 统计学 2024-07-08 Jakob Raymaekers , Peter J. Rousseeuw

The conditional value-at-risk (CVaR) is a useful risk measure in fields such as machine learning, finance, insurance, energy, etc. When measuring very extreme risk, the commonly used CVaR estimation method of sample averaging does not work…

统计方法学 · 统计学 2021-03-10 Dylan Troop , Frédéric Godin , Jia Yuan Yu

The term moderate deviations is often used in the literature to mean a class of large deviation principles that, in some sense, fills the gap between a convergence in probability of some random variables to a constant and a weak convergence…

概率论 · 数学 2024-11-20 Rita Giuliano , Claudio Macci , Barbara Pacchiarotti

A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label…

机器学习 · 统计学 2021-11-17 Nilesh Tripuraneni , Ben Adlam , Jeffrey Pennington

Several new geometric quantile-based measures for multivariate dispersion, skewness, kurtosis, and spherical asymmetry are defined. These measures differ from existing measures, which use volumes and are easy to calculate. Some theoretical…

统计理论 · 数学 2024-12-30 Ha-Young Shin , Hee-Seok Oh

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so…

计算与语言 · 计算机科学 2018-06-20 Henry B. Moss , David S. Leslie , Paul Rayson

Cross-validation (CV) is a common method to tune machine learning methods and can be used for model selection in regression as well. Because of the structured nature of small, traditional experimental designs, the literature has warned…

应用统计 · 统计学 2025-06-18 Maria L. Weese , Byran J. Smucker , David J. Edwards

We consider the problem of estimating the mean of a random vector based on i.i.d. observations and adversarial contamination. We introduce a multivariate extension of the trimmed-mean estimator and show its optimal performance under minimal…

统计理论 · 数学 2020-02-25 Gabor Lugosi , Shahar Mendelson

Cross-validation (CV) is a technique for evaluating the ability of statistical models/learning systems based on a given data set. Despite its wide applicability, the rather heavy computational cost can prevent its use as the system size…

机器学习 · 统计学 2016-10-26 Yoshiyuki Kabashima , Tomoyuki Obuchi , Makoto Uemura

In Bayesian analysis, the posterior follows from the data and a choice of a prior and a likelihood. One hopes that the posterior is robust to reasonable variation in the choice of prior and likelihood, since this choice is made by the…

统计方法学 · 统计学 2015-12-09 Ryan Giordano , Tamara Broderick , Michael Jordan