English
Related papers

Related papers: Robust regression with covariate filtering: Heavy …

200 papers

The composite quantile regression (CQR) was introduced by Zou and Yuan [Ann. Statist. 36 (2008) 1108--1126] as a robust regression method for linear models with heavy-tailed errors while achieving high efficiency. Its penalized counterpart…

Methodology · Statistics 2023-10-16 Haeseong Moon , Wen-Xin Zhou

This paper investigates the stability of deep ReLU neural networks for nonparametric regression under the assumption that the noise has only a finite p-th moment. We unveil how the optimal rate of convergence depends on p, the degree of…

Statistics Theory · Mathematics 2023-01-02 Jianqing Fan , Yihong Gu , Wen-Xin Zhou

Recently, high-dimensional heterogeneous data have attracted a lot of attention and discussion. Under heterogeneity, semiparametric regression is a popular choice to model data in statistics. In this paper, we take advantages of expectile…

Statistics Theory · Mathematics 2019-08-20 Jun Zhao , Guan'ao Yan , Yi Zhang

The problem of estimating the coefficient of bivariate tail dependence is considered here from the robustness point of view; it combines two apparently contradictory theories of robust statistics and extreme value statistics. The usual…

Applications · Statistics 2014-07-08 Abhik Ghosh

The least squares of depth-trimmed (LST) residuals regression, proposed and studied in Zuo and Zuo (2023), serves as a robust alternative to the classic least squares (LS) regression as well as a strong competitor to the renowned robust…

Applications · Statistics 2025-01-28 Yijun Zuo , Hanwen Zuo

In the famous least sum of trimmed squares (LTS) of residuals estimator (Rousseeuw (1984)), residuals are first squared and then trimmed. In this article, we first trim residuals - using a depth trimming scheme - and then square the rest of…

Methodology · Statistics 2022-11-28 Yijun Zuo

We consider the problem of estimating the mean of a random vector based on i.i.d. observations and adversarial contamination. We introduce a multivariate extension of the trimmed-mean estimator and show its optimal performance under minimal…

Statistics Theory · Mathematics 2020-02-25 Gabor Lugosi , Shahar Mendelson

Large datasets are often affected by cell-wise outliers in the form of missing or erroneous data. However, discarding any samples containing outliers may result in a dataset that is too small to accurately estimate the covariance matrix.…

Statistics Theory · Mathematics 2023-11-13 Karim Lounici , Grégoire Pacreau

We consider the question of Gaussian mean testing, a fundamental task in high-dimensional distribution testing and signal processing, subject to adversarial corruptions of the samples. We focus on the relative power of different…

Data Structures and Algorithms · Computer Science 2023-07-21 Clément L. Canonne , Samuel B. Hopkins , Jerry Li , Allen Liu , Shyam Narayanan

In the classical contamination models, such as the gross-error (Huber and Tukey contamination model or Case-wise Contamination), observations are considered as the units to be identified as outliers or not. This model is very useful when…

Statistics Theory · Mathematics 2021-03-11 Giovanni Saraceno , Claudio Agostinelli

We propose a robust Bayesian formulation of random feature (RF) regression that accounts explicitly for prior and likelihood misspecification via Huber-style contamination sets. Starting from the classical equivalence between…

Machine Learning · Computer Science 2026-02-24 Michele Caprio , Katerina Papagiannouli , Siu Lun Chau , Sayan Mukherjee

Datasets with extreme observations and/or heavy-tailed error distributions are commonly encountered and should be analyzed with careful consideration of these features from a statistical perspective. Small deviations from an assumed model,…

Methodology · Statistics 2023-01-12 Meadhbh O'Neill , Kevin Burke

Due to the highly non-convex nature of large-scale robust parameter estimation, avoiding poor local minima is challenging in real-world applications where input data is contaminated by a large or unknown fraction of outliers. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2020-03-23 Huu Le , Christopher Zach

In this paper, we investigate the adversarial robustness of nonparametric regression, a fundamental problem in machine learning, under the setting where an adversary can arbitrarily corrupt a subset of the input data. While the robustness…

Machine Learning · Computer Science 2025-10-28 Parsa Moradi , Hanzaleh Akabrinodehi , Mohammad Ali Maddah-Ali

This paper tackles the problem of robust covariance matrix estimation when the data is incomplete. Classical statistical estimation methodologies are usually built upon the Gaussian assumption, whereas existing robust estimation ones assume…

The Seemingly Unrelated Regressions (SUR) model is a wide used estimation procedure in econometrics, insurance and finance, where very often, the regression model contains more than one equation. Unknown parameters, regression coefficients…

Methodology · Statistics 2021-07-05 Giovanni Saraceno , Fatemah Alqallaf , Claudio Agostinelli

Confounding variables are a recurrent challenge for causal discovery and inference. In many situations, complex causal mechanisms only manifest themselves in extreme events, or take simpler forms in the extremes. Stimulated by data on…

Methodology · Statistics 2024-11-14 Olivier C. Pasche , Valérie Chavez-Demoulin , Anthony C. Davison

We study the problem of computationally efficient robust estimation of the covariance/scatter matrix of elliptical distributions -- that is, affine transformations of spherically symmetric distributions -- under the strong contamination…

Data Structures and Algorithms · Computer Science 2025-04-15 Gleb Novikov

Low-rank tensor models are widely used in statistics. However, most existing methods rely heavily on the assumption that data follows a sub-Gaussian distribution. To address the challenges associated with heavy-tailed distributions…

Methodology · Statistics 2025-09-16 Xiaoyu Zhang , Di Wang , Guodong Li , Defeng Sun

We study fast algorithms for statistical regression problems under the strong contamination model, where the goal is to approximately optimize a generalized linear model (GLM) given adversarially corrupted samples. Prior works in this line…

Data Structures and Algorithms · Computer Science 2021-06-23 Arun Jambulapati , Jerry Li , Tselil Schramm , Kevin Tian