English
Related papers

Related papers: A Unified Approach to Robust Mean Estimation

200 papers

Learning in the presence of outliers is a fundamental problem in statistics. Until recently, all known efficient unsupervised learning algorithms were very sensitive to outliers in high dimensions. In particular, even for the task of robust…

Data Structures and Algorithms · Computer Science 2019-11-15 Ilias Diakonikolas , Daniel M. Kane

The Median of Means (MoM) is a mean estimator that has gained popularity in the context of heavy-tailed data. In this work, we analyze its performance in the task of simultaneously estimating the mean of each function in a class…

Machine Learning · Statistics 2025-06-23 Mikael Møller Høgsgaard , Andrea Paudice

Meta-analyses frequently include trials that report multiple effect sizes based on a common set of study participants. These effect sizes will generally be correlated. Cluster-robust variance-covariance estimators are a fruitful approach…

Methodology · Statistics 2022-03-07 Thilo Welz , Wolfgang Viechtbauer , Markus Pauly

While there is a rich literature on robust methodologies for contamination in continuously distributed data, contamination in categorical data is largely overlooked. This is regrettable because many datasets are categorical and oftentimes…

Methodology · Statistics 2024-12-13 Max Welz

It is of importance to develop statistical techniques to analyze high-dimensional data in the presence of both complex dependence and possible outliers in real-world applications such as imaging data analyses. We propose a new robust…

Methodology · Statistics 2021-10-01 Bingyuan Liu , Qi Zhang , Lingzhou Xue , Peter X. K. Song , Jian Kang

Tracking underwater autonomous platforms is often difficult because of noisy, biased, and discretized input data. Classic filters and smoothers based on standard assumptions of Gaussian white noise break down when presented with any of…

Optimization and Control · Mathematics 2019-05-24 Jonathan Jonker , Aleksandr Aravkin , James V. Burke , Gianluigi Pillonetto , Sarah Webster

In this paper, we study robust covariance estimation under the approximate factor model with observed factors. We propose a novel framework to first estimate the initial joint covariance matrix of the observed data and the factors, and then…

Methodology · Statistics 2016-02-03 Jianqing Fan , Weichen Wang , Yiqiao Zhong

We study the problem of estimating a $p$-dimensional $s$-sparse vector in a linear model with Gaussian design and additive noise. In the case where the labels are contaminated by at most $o$ adversarial outliers, we prove that the…

Statistics Theory · Mathematics 2019-11-20 Arnak S. Dalalyan , Philip Thompson

Complex distributions of the healthcare expenditure pose challenges to statistical modeling via a single model. Super learning, an ensemble method that combines a range of candidate models, is a promising alternative for cost estimation and…

Machine Learning · Statistics 2022-05-17 Ziyue Wu , David Benkeser

Quantiles and expectiles are determined by different loss functions: asymmetric least absolute deviation for quantiles and asymmetric squared loss for expectiles. This distinction ensures that quantile regression methods are robust to…

Methodology · Statistics 2025-10-08 Abdellah Atanane , Abdallah Mkhadri , Karim Oualkacha

Robust estimation of a mean vector, a topic regarded as obsolete in the traditional robust statistics community, has recently surged in machine learning literature in the last decade. The latest focus is on the sub-Gaussian performance and…

Machine Learning · Statistics 2022-02-22 Yijun Zuo

This paper investigates tradeoffs among optimization errors, statistical rates of convergence and the effect of heavy-tailed errors for high-dimensional robust regression with nonconvex regularization. When the additive errors in linear…

Statistics Theory · Mathematics 2021-01-01 Xiaoou Pan , Qiang Sun , Wen-Xin Zhou

We consider the task of heavy-tailed statistical estimation given streaming $p$-dimensional samples. This could also be viewed as stochastic optimization under heavy-tailed distributions, with an additional $O(p)$ space complexity…

Machine Learning · Computer Science 2022-02-28 Che-Ping Tsai , Adarsh Prasad , Sivaraman Balakrishnan , Pradeep Ravikumar

This paper addresses the scalar regression problem through a novel solution to exactly optimize the Huber loss in a general semi-supervised setting, which combines multi-view learning and manifold regularization. We propose a principled…

Machine Learning · Computer Science 2016-06-28 Jacopo Cavazza , Vittorio Murino

We investigate robustness properties of pre-trained neural models for automatic speech recognition. Real life data in machine learning is usually very noisy and almost never clean, which can be attributed to various factors depending on the…

Computation and Language · Computer Science 2022-08-19 Goutham Rajendran , Wei Zou

In existing distributed stochastic optimization studies, it is usually assumed that the gradient noise has a bounded variance. However, recent research shows that the heavy-tailed noise, which allows an unbounded variance, is closer to…

Optimization and Control · Mathematics 2025-05-15 Jun Hu , Chao Sun , Bo Chen , Jianzheng Wang , Zheming Wang

Semi-functional linear regression models postulate a linear relationship between a scalar response and a functional covariate, and also include a non-parametric component involving a univariate explanatory variable. It is of practical…

Methodology · Statistics 2023-08-08 Graciela Boente , Matias Salibian-Barrera , Pablo Vena

In this paper, we study the problem of sparse mean estimation under adversarial corruptions, where the goal is to estimate the $k$-sparse mean of a heavy-tailed distribution from samples contaminated by adversarial noise. Existing methods…

Machine Learning · Computer Science 2025-08-26 Jianhao Ma , Rui Ray Chen , Yinghui He , Salar Fattahi , Wei Hu

A robust summarization system should be able to capture the gist of the document, regardless of the specific word choices or noise in the input. In this work, we first explore the summarization models' robustness against perturbations…

Computation and Language · Computer Science 2023-06-05 Xiuying Chen , Guodong Long , Chongyang Tao , Mingzhe Li , Xin Gao , Chengqi Zhang , Xiangliang Zhang

Information from various data sources is increasingly available nowadays. However, some of the data sources may produce biased estimation due to commonly encountered biased sampling, population heterogeneity, or model misspecification. This…

Methodology · Statistics 2023-02-07 Ruoyu Wang , Qihua Wang , Wang Miao