English
Related papers

Related papers: Provable More Data Hurt in High Dimensional Least …

200 papers

The "double descent" risk curve was proposed to qualitatively describe the out-of-sample prediction accuracy of variably-parameterized machine learning models. This article provides a precise mathematical analysis for the shape of this…

Machine Learning · Computer Science 2020-12-22 Mikhail Belkin , Daniel Hsu , Ji Xu

This study proposes a point estimator of the break location for a one-time structural break in linear regression models. If the break magnitude is small, the least-squares estimator of the break date has two modes at the ends of the finite…

Econometrics · Economics 2020-06-04 Yaein Baek

In this article, we study the limit distribution of the least square estimator, properly normalized, from a regression model in which observations are assumed to be finite ($\alpha N$) and sampled under two different random times. Based on…

Statistics Theory · Mathematics 2020-12-17 Tania Roa , Soledad Torres , Ciprian tudor

We investigate the theoretical performances of the Partial Least Square (PLS) algorithm in a high dimensional context. We provide upper bounds on the risk in prediction for the statistical linear model when considering the PLS estimator.…

Statistics Theory · Mathematics 2024-10-15 Luca Castelli , Irène Gannaz , Clément Marteau

This paper focuses on the problem of the estimation of the cumulative hazard function of a distribution on a general complete separable metric space when the data points are subject to censoring by an arbitrary adapted random set. A problem…

Statistics Theory · Mathematics 2013-09-04 Alberto Carabarin Aguirre , B. Gail Ivanoff

Good robust estimators can be tuned to combine a high breakdown point and a specified asymptotic efficiency at a central model. This happens in regression with MM- and tau-estimators among others. However, the finite-sample efficiency of…

Statistics Theory · Mathematics 2013-11-21 Ricardo Maronna , Víctor Yohai

This paper is concerned with the problem of policy evaluation with linear function approximation in discounted infinite horizon Markov decision processes. We investigate the sample complexities required to guarantee a predefined estimation…

Machine Learning · Statistics 2024-05-03 Gen Li , Weichen Wu , Yuejie Chi , Cong Ma , Alessandro Rinaldo , Yuting Wei

This paper proposes a max-test for testing (possibly infinitely) many zero parameter restrictions in an extremum estimation framework. The test statistic is formed by estimating key parameters one at a time based on many empirical loss…

Statistics Theory · Mathematics 2022-04-12 Jonathan B. Hill

We prove a new and general concentration inequality for the excess risk in least-squares regression with random design and heteroscedastic noise. No specific structure is required on the model, except the existence of a suitable function…

Statistics Theory · Mathematics 2018-03-12 Adrien Saumard

In this paper, under mild assumptions, we derive a law of large numbers, a central limit theorem with an error estimate, an almost sure invariance principle and a variant of Chernoff bound in finite-state hidden Markov models. These limit…

Information Theory · Computer Science 2012-04-13 Guangyue Han

Overparametrization often helps improve the generalization performance. This paper presents a dual view of overparametrization suggesting that downsampling may also help generalize. Focusing on the proportional regime $m\asymp n \asymp p$,…

Statistics Theory · Mathematics 2023-10-17 Xin Chen , Yicheng Zeng , Siyue Yang , Qiang Sun

The purpose of this article is to develop a general parametric estimation theory that allows the derivation of the limit distribution of estimators in non-regular models where the true parameter value may lie on the boundary of the…

Statistics Theory · Mathematics 2022-11-28 Junichiro Yoshida , Nakahiro Yoshida

We propose new model selection criteria based on generalized ridge estimators dominating the maximum likelihood estimator under the squared risk and the Kullback-Leibler risk in multivariate linear regression. Our model selection criteria…

Statistics Theory · Mathematics 2016-04-08 Yuichi Mori , Taiji Suzuki

Error-in-variables regression is a common ingredient in treatment effect estimators using panel data. This includes synthetic control estimators, counterfactual time series forecasting estimators, and combinations. We study high-dimensional…

Statistics Theory · Mathematics 2021-04-20 David A. Hirshberg

We address the problem of estimating the expected shortfall risk of a financial loss using a finite number of i.i.d. data. It is well known that the classical plug-in estimator suffers from poor statistical performance when faced with…

Risk Management · Quantitative Finance 2026-02-13 Daniel Bartl , Stephan Eckstein

For Huber contamination on a known finite sample space, the unrestricted contaminating law is a probability vector on the support atoms, and domination over all measurable subsets reduces to atomwise inequalities. Placing a Dirichlet prior…

Methodology · Statistics 2026-05-27 Jaehoan Kim

Missing values arise in most real-world data sets due to the aggregation of multiple sources and intrinsically missing information (sensor failure, unanswered questions in surveys...). In fact, the very nature of missing values usually…

Machine Learning · Statistics 2022-02-04 Alexis Ayme , Claire Boyer , Aymeric Dieuleveut , Erwan Scornet

We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…

Machine Learning · Computer Science 2024-12-03 Maryam Aliakbarpour , Piotr Indyk , Ronitt Rubinfeld , Sandeep Silwal

To take sample biases and skewness in the observations into account, practitioners frequently weight their observations according to some marginal distribution. The present paper demonstrates that such weighting can indeed improve the…

Methodology · Statistics 2018-11-05 Tobias Niebuhr , Mathias Trabs

This paper extends the standard chaining technique to prove excess risk upper bounds for empirical risk minimization with random design settings even if the magnitude of the noise and the estimates is unbounded. The bound applies to many…

Machine Learning · Statistics 2016-09-08 Gábor Balázs , András György , Csaba Szepesvári