English
Related papers

Related papers: Provable More Data Hurt in High Dimensional Least …

200 papers

We study the excess minimum risk in statistical inference, defined as the difference between the minimum expected loss in estimating a random variable from an observed feature vector and the minimum expected loss in estimating the same…

Information Theory · Computer Science 2023-09-29 László Györfi , Tamás Linder , Harro Walk

Motivated by the prevalence of environments in which data is abundant while resources for storage and/or transmission might be scarce, we study linear regression when predictors, their squares, and responses are subject to single-bit…

Statistics Theory · Mathematics 2026-04-01 Daniel Hill , Martin Slawski

In this short note, we provide a sample complexity lower bound for learning linear predictors with respect to the squared loss. Our focus is on an agnostic setting, where no assumptions are made on the data distribution. This contrasts with…

Machine Learning · Computer Science 2021-11-23 Ohad Shamir

The statistics and machine learning communities have recently seen a growing interest in classification-based approaches to two-sample testing. The outcome of a classification-based two-sample test remains a rejection decision, which is not…

Statistics Theory · Mathematics 2022-11-15 Loris Michel , Jeffrey Näf , Nicolai Meinshausen

This manuscript studies statistical properties of linear classifiers obtained through minimization of an unregularized convex risk over a finite sample. Although the results are explicitly finite-dimensional, inputs may be passed through…

Machine Learning · Computer Science 2012-06-15 Matus Telgarsky

When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and…

Methodology · Statistics 2020-08-11 Chanseok Park , Haewon Kim , Min Wang

Linear Least Squares is a very well known technique for parameter estimation, which is used even when sub-optimal, because of its very low computational requirements and the fact that exact knowledge of the noise statistics is not required.…

Signal Processing · Electrical Eng. & Systems 2017-11-01 Michael Krikheli , Amir Leshem

Scientific explanation often requires inferring maximally predictive features from a given data set. Unfortunately, the collection of minimal maximally predictive features for most stochastic processes is uncountably infinite. In such…

Statistical Mechanics · Physics 2017-05-31 Sarah E. Marzen , James P. Crutchfield

In recent years, there has been a growing interest in the effects of data poisoning attacks on data-driven control methods. Poisoning attacks are well-known to the Machine Learning community, which, however, make use of assumptions, such as…

Systems and Control · Electrical Eng. & Systems 2023-05-17 Alessio Russo

In this work, we focus on the high-dimensional trace regression model with a low-rank coefficient matrix. We establish a nearly optimal in-sample prediction risk bound for the rank-constrained least-squares estimator under no assumptions on…

Statistics Theory · Mathematics 2022-04-19 Michael Law , Ya'acov Ritov , Ruixiang Zhang , Ziwei Zhu

Robust statistics aims to compute quantities to represent data where a fraction of it may be arbitrarily corrupted. The most essential statistic is the mean, and in recent years, there has been a flurry of theoretical advancement for…

Machine Learning · Statistics 2025-02-18 Cullen Anderson , Jeff M. Phillips

We study the statistical properties of the least squares estimator in unimodal sequence estimation. Although closely related to isotonic regression, unimodal regression has not been as extensively studied. We show that the unimodal least…

Statistics Theory · Mathematics 2017-05-10 Sabyasachi Chatterjee , John Lafferty

The least trimmed squares (LTS) estimator is a renowned robust alternative to the classic least squares estimator and is popular in location, regression, machine learning, and AI literature. Many studies exist on LTS, including its…

Machine Learning · Statistics 2025-01-10 Yijun Zuo

We consider least squares estimation in a general nonparametric regression model. The rate of convergence of the least squares estimator (LSE) for the unknown regression function is well studied when the errors are sub-Gaussian. We find…

Statistics Theory · Mathematics 2021-04-12 Arun K. Kuchibhotla , Rohit K. Patra

This paper is concerned with estimation and inference for ultrahigh dimensional partially linear single-index models. The presence of high dimensional nuisance parameter and nuisance unknown function makes the estimation and inference…

Methodology · Statistics 2024-04-09 Shijie Cui , Xu Guo , Zhe Zhang

Given a sample of i.i.d. high-dimensional centered random vectors, we consider a problem of estimation of their covariance matrix $\Sigma$ with an additional assumption that $\Sigma$ can be represented as a sum of a few Kronecker products…

Statistics Theory · Mathematics 2024-06-18 Nikita Puchkin , Maxim Rakhuba

We consider the high-dimensional inference problem where the signal is a low-rank matrix which is corrupted by an additive Gaussian noise. Given a probabilistic model for the low-rank matrix, we compute the limit in the large dimension…

Probability · Mathematics 2018-06-01 Léo Miolane

A continuous-time regression model with a jointly strictly sub-Gaussian random noise is considered in the paper. Upper exponential bounds for probabilities of large deviations of the least squares estimator for the regression parameter are…

Probability · Mathematics 2018-06-12 Alexander V. Ivanov , Igor V. Orlovskyi

We consider the problem of fitting the parameters of a high-dimensional linear regression model. In the regime where the number of parameters $p$ is comparable to or exceeds the sample size $n$, a successful approach uses an…

Statistics Theory · Mathematics 2013-11-04 Adel Javanmard , Andrea Montanari

We analyze the problem of discrete distribution estimation under $\ell_1$ loss. We provide non-asymptotic upper and lower bounds on the maximum risk of the empirical distribution (the maximum likelihood estimator), and the minimax risk in…

Information Theory · Computer Science 2015-12-31 Yanjun Han , Jiantao Jiao , Tsachy Weissman