English
Related papers

Related papers: Generalized equivalences between subsampling and r…

200 papers

Consider estimation of the regression function based on a model with equidistant design and measurement errors generated from a fractional Gaussian noise process. In previous literature, this model has been heuristically linked to an…

Statistics Theory · Mathematics 2014-12-02 Johannes Schmidt-Hieber

This paper introduces a flexible regularization approach that reduces point estimation risk of group means stemming from e.g. categorical regressors, (quasi-)experimental data or panel data models. The loss function is penalized by adding…

Econometrics · Economics 2019-01-08 Phillip Heiler , Jana Mareckova

Kernel methods for deconvolution have attractive features, and prevail in the literature. However, they have disadvantages, which include the fact that they are usually suitable only for cases where the error distribution is infinitely…

Statistics Theory · Mathematics 2009-09-29 Peter Hall , Alexander Meister

We prove a non-asymptotic distribution-independent lower bound for the expected mean squared generalization error caused by label noise in ridgeless linear regression. Our lower bound generalizes a similar known result to the…

Machine Learning · Statistics 2023-08-02 David Holzmüller

High-dimensional data analysis has motivated a spectrum of regularization methods for variable selection and sparse modeling, with two popular classes of convex ones and concave ones. A long debate has been on whether one class dominates…

Methodology · Statistics 2016-05-12 Yingying Fan , Jinchi Lv

Regression trees and random forests are popular and effective non-parametric estimators in practical applications. A recent paper by Athey and Wager shows that the random forest estimate at any point is asymptotically Gaussian; in this…

Econometrics · Economics 2021-02-02 Kevin Li

Feature bagging is a well-established ensembling method which aims to reduce prediction variance by combining predictions of many estimators trained on subsets or projections of features. Here, we develop a theory of feature-bagging in…

Machine Learning · Statistics 2024-01-11 Benjamin S. Ruben , Cengiz Pehlevan

This work unifies the analysis of various randomized methods for solving linear and nonlinear inverse problems by framing the problem in a stochastic optimization setting. By doing so, we show that many randomized methods are variants of a…

Numerical Analysis · Mathematics 2023-06-21 Jonathan Wittmer , C. G. Krishnanunni , Hai V. Nguyen , Tan Bui-Thanh

We study properties of ridge functions $f(x)=g(a\cdot x)$ in high dimensions $d$ from the viewpoint of approximation theory. The considered function classes consist of ridge functions such that the profile $g$ is a member of a univariate…

Numerical Analysis · Mathematics 2013-11-11 Sebastian Mayer , Tino Ullrich , Jan Vybiral

We consider stochastic optimization problems which use observed data to estimate essential characteristics of the random quantities involved. Sample average approximation (SAA) or empirical (plug-in) estimation are very popular ways to use…

Statistics Theory · Mathematics 2021-03-16 Darinka Dentcheva , Yang Lin

For finite samples with binary outcomes penalized logistic regression such as ridge logistic regression (RR) has the potential of achieving smaller mean squared errors (MSE) of coefficients and predictions than maximum likelihood…

Methodology · Statistics 2021-01-28 Hana Šinkovec , Georg Heinze , Rok Blagus , Angelika Geroldinger

We provide a statistical analysis of regularization-based continual learning on a sequence of linear regression tasks, with emphasis on how different regularization terms affect the model performance. We first derive the convergence rate…

Machine Learning · Computer Science 2024-06-11 Xuyang Zhao , Huiyuan Wang , Weiran Huang , Wei Lin

We introduce the concept of coverage risk as an error measure for density ridge estimation. The coverage risk generalizes the mean integrated square error to set estimation. We propose two risk estimators for the coverage risk and we show…

Methodology · Statistics 2015-06-09 Yen-Chi Chen , Christopher R. Genovese , Shirley Ho , Larry Wasserman

We study the monotone single index model where a real response variable $Y $ is linked to a $d$-dimensional covariate $X$ through the relationship $E[Y | X] = \Psi_0(\alpha^T_0 X)$ almost surely. Both the ridge function, $\Psi_0$, and the…

Statistics Theory · Mathematics 2018-04-19 F. Balabdaoui , C. Durot , H. Jankowski

Ridge regression (RR) is an important machine learning technique which introduces a regularization hyperparameter $\alpha$ to ordinary multiple linear regression for analyzing data suffering from multicollinearity. In this paper, we present…

Quantum Physics · Physics 2021-08-03 Chao-Hua Yu , Fei Gao , Qiao-Yan Wen

Many standard estimators, when applied to adaptively collected data, fail to be asymptotically normal, thereby complicating the construction of confidence intervals. We address this challenge in a semi-parametric context: estimating the…

Statistics Theory · Mathematics 2025-03-04 Licong Lin , Koulik Khamaru , Martin J. Wainwright

This article provides, through theoretical analysis, an in-depth understanding of the classification performance of the empirical risk minimization framework, in both ridge-regularized and unregularized cases, when high dimensional data are…

Machine Learning · Statistics 2020-11-26 Xiaoyi Mai , Zhenyu Liao

The ever-growing size of the datasets renders well-studied learning techniques, such as Kernel Ridge Regression, inapplicable, posing a serious computational challenge. Divide-and-conquer is a common remedy, suggesting to split the dataset…

Machine Learning · Statistics 2021-05-25 Valeriy Avanesov

Estimating linear, mean-square continuous functionals is a pivotal challenge in statistics. In high-dimensional contexts, this estimation is often performed under the assumption of exact model sparsity, meaning that only a small number of…

Statistics Theory · Mathematics 2025-08-04 Jelena Bradic , Victor Chernozhukov , Whitney K. Newey , Yinchu Zhu

In high dimension, it is customary to consider Lasso-type estimators to enforce sparsity. For standard Lasso theory to hold, the regularization parameter should be proportional to the noise level, yet the latter is generally unknown in…

Machine Learning · Statistics 2017-10-19 Mathurin Massias , Olivier Fercoq , Alexandre Gramfort , Joseph Salmon