English
Related papers

Related papers: Moving Beyond Sub-Gaussianity in High-Dimensional …

200 papers

We propose and analyze a new estimator of the covariance matrix that admits strong theoretical guarantees under weak assumptions on the underlying distribution, such as existence of moments of only low order. While estimation of covariance…

Statistics Theory · Mathematics 2018-01-17 Stanislav Minsker , Xiaohan Wei

This paper studies the Gaussian approximation of high-dimensional and non-degenerate U-statistics of order two under the supremum norm. We propose a two-step Gaussian approximation procedure that does not impose structural assumptions on…

Statistics Theory · Mathematics 2016-10-04 Xiaohui Chen

We consider least squares estimation in a general nonparametric regression model. The rate of convergence of the least squares estimator (LSE) for the unknown regression function is well studied when the errors are sub-Gaussian. We find…

Statistics Theory · Mathematics 2021-04-12 Arun K. Kuchibhotla , Rohit K. Patra

The Lasso is one of the most important approaches for parameter estimation and variable selection in high dimensional linear regression. At the heart of its success is the attractive rate of convergence result even when $p$, the dimension…

Statistics Theory · Mathematics 2019-08-09 Junlong Zhao , Chenlei Leng

We propose a residual randomization procedure designed for robust Lasso-based inference in the high-dimensional setting. Compared to earlier work that focuses on sub-Gaussian errors, the proposed procedure is designed to work robustly in…

Methodology · Statistics 2021-08-20 Y. Samuel Wang , Si Kai Lee , Panos Toulis , Mladen Kolar

``Localization'' has proven to be a valuable tool in the Statistical Learning literature as it allows sharp risk bounds in terms of the problem geometry. Localized bounds seem to be much less exploited in the Stochastic Optimization…

Optimization and Control · Mathematics 2023-03-30 Roberto I. Oliveira , Philip Thompson

In this paper, we propose a uniformly dithered 1-bit quantization scheme for high-dimensional statistical estimation. The scheme contains truncation, dithering, and quantization as typical steps. As canonical examples, the quantization…

Machine Learning · Statistics 2023-01-23 Junren Chen , Cheng-Long Wang , Michael K. Ng , Di Wang

Stochastic gradient descent is one of the most common iterative algorithms used in machine learning and its convergence analysis is a rich area of research. Understanding its convergence properties can help inform what modifications of it…

Optimization and Control · Mathematics 2025-11-25 Liam Madden , Emiliano Dall'Anese , Stephen Becker

Sums of independent, bounded random variables concentrate around their expectation approximately as well a Gaussian of the same variance. Well known results of this form include the Bernstein, Hoeffding, and Chernoff inequalities and many…

Discrete Mathematics · Computer Science 2017-04-25 Thomas Steinke , Jonathan Ullman

We propose the notion of sub-Weibull distributions, which are characterised by tails lighter than (or equally light as) the right tail of a Weibull distribution. This novel class generalises the sub-Gaussian and sub-Exponential families to…

Statistics Theory · Mathematics 2020-12-04 Mariia Vladimirova , Stephane Girard , Hien Nguyen , Julyan Arbel

Datasets with extreme observations and/or heavy-tailed error distributions are commonly encountered and should be analyzed with careful consideration of these features from a statistical perspective. Small deviations from an assumed model,…

Methodology · Statistics 2023-01-12 Meadhbh O'Neill , Kevin Burke

Expected shortfall is defined as the average over the tail below (or above) a certain quantile of a probability distribution. Expected shortfall regression provides powerful tools for learning the relationship between a response variable…

Methodology · Statistics 2025-01-03 Shushu Zhang , Xuming He , Kean Ming Tan , Wen-Xin Zhou

High-dimensional data arise routinely in modern statistics, econometrics, finance, genomics, and machine learning. While a large body of existing methodology is developed under Gaussian or light-tailed assumptions, many real data sets…

Methodology · Statistics 2026-04-16 Long Feng

Stochastic iterative methods are useful in a variety of large-scale numerical linear algebraic, machine learning, and statistical problems, in part due to their low-memory footprint. They are frequently used in a variety of applications,…

Numerical Analysis · Mathematics 2025-11-27 Toby Anderson , Max Collins , Jamie Haddock , Jackie Lok , Elizaveta Rebrova

Our goal is to develop a Bayesian model averaging technique in linear regression models that accommodates heavier tailed error densities than the normal distribution. Motivated by the use of the Huber loss function in the presence of…

Methodology · Statistics 2024-11-26 Shamriddha De , Joyee Ghosh

A current strand of research in high-dimensional statistics deals with robustifying the available methodology with respect to deviations from the pervasive light-tail assumptions. In this paper we consider a linear mean regression model…

Statistics Theory · Mathematics 2025-02-06 Philipp Hermann , Hajo Holzmann

Recent studies have shown that heavy tails can emerge in stochastic optimization and that the heaviness of the tails have links to the generalization error. While these studies have shed light on interesting aspects of the generalization…

Machine Learning · Statistics 2023-02-14 Anant Raj , Melih Barsbey , Mert Gürbüzbalaban , Lingjiong Zhu , Umut Şimşekli

Models for extreme values are generally derived from limit results, which are meant to be good enough approximations when applied to finite samples. Depending on the speed of convergence of the process underlying the data, these…

Statistics Theory · Mathematics 2019-02-20 Thomas Lugrin , Anthony C. Davison , Jonathan A. Tawn

We propose a self-tuning $\sqrt{\mathrm {Lasso}}$ method that simultaneously resolves three important practical problems in high-dimensional regression analysis, namely it handles the unknown scale, heteroscedasticity and (drastic)…

Methodology · Statistics 2014-05-27 Alexandre Belloni , Victor Chernozhukov , Lie Wang

We study the algorithmic problem of estimating the mean of heavy-tailed random vector in $\mathbb{R}^d$, given $n$ i.i.d. samples. The goal is to design an efficient estimator that attains the optimal sub-gaussian error bound, only assuming…

Statistics Theory · Mathematics 2020-02-19 Zhixian Lei , Kyle Luh , Prayaag Venkat , Fred Zhang