English
Related papers

Related papers: Bagging in overparameterized learning: Risk charac…

200 papers

Cross-validation is a well-known and widely used bandwidth selection method in nonparametric regression estimation. However, this technique has two remarkable drawbacks: (i) the large variability of the selected bandwidths, and (ii) the…

Methodology · Statistics 2021-05-11 D. Barreiro-Ures , R. Cao , M. Francisco-Fernández

Interpolators are unstable. For example, the mininum $\ell_2$ norm least square interpolator exhibits unbounded test errors when dealing with noisy data. In this paper, we study how ensemble stabilizes and thus improves the generalization…

Machine Learning · Statistics 2023-09-08 Mingqi Wu , Qiang Sun

Hall and Robinson (2009) proposed and analyzed the use of bagged cross-validation to choose the bandwidth of a kernel density estimator. They established that bagging greatly reduces the noise inherent in ordinary cross-validation, and…

Methodology · Statistics 2024-02-01 Daniel Barreiro-Ures , Ricardo Cao , Mario Francisco Fernández , Jeffrey D. Hart

We consider the application of a popular penalised regression method, Ridge Regression, to data with very high dimensions and many more covariates than observations. Our motivation is the problem of out-of-sample prediction and the setting…

Applications · Statistics 2012-05-04 Erika Cule , Maria De Iorio

Machine learning techniques always aim to reduce the generalized prediction error. In order to reduce it, ensemble methods present a good approach combining several models that results in a greater forecasting capacity. The Random Machines…

Machine Learning · Statistics 2020-03-31 Anderson Ara , Mateus Maia , Samuel Macêdo , Francisco Louzada

Feature bagging is a well-established ensembling method which aims to reduce prediction variance by combining predictions of many estimators trained on subsets or projections of features. Here, we develop a theory of feature-bagging in…

Machine Learning · Statistics 2024-01-11 Benjamin S. Ruben , Cengiz Pehlevan

We explore the performance of sample average approximation in comparison with several other methods for stochastic optimization when there is information available on the underlying true probability distribution. The methods we evaluate are…

Machine Learning · Computer Science 2019-07-22 Eddie Anderson , Harrison Nguyen

When randomized ensembles such as bagging or random forests are used for binary classification, the prediction error of the ensemble tends to decrease and stabilize as the number of classifiers increases. However, the precise relationship…

Probability · Mathematics 2019-05-01 Miles E. Lopes

When randomized ensemble methods such as bagging and random forests are implemented, a basic question arises: Is the ensemble large enough? In particular, the practitioner desires a rigorous guarantee that a given ensemble will perform…

Machine Learning · Statistics 2019-08-06 Miles E. Lopes , Suofei Wu , Thomas C. M. Lee

Subsampling is a popular approach to alleviating the computational burden for analyzing massive datasets. Recent efforts have been devoted to various statistical models without explicit regularization. In this paper, we develop an efficient…

Methodology · Statistics 2022-04-12 Yunlu Chen , Nan Zhang

For a larger set of predictions of several differently trained machine learning models, known as bagging predictors, the mean of all predictions is taken by default. Nevertheless, this proceeding can deviate from the actual ground truth in…

Machine Learning · Computer Science 2026-04-07 Philipp Seitz , Jan Schmitt , Andreas Schiffler

Predictions from machine learning algorithms can vary across random seeds, inducing instability in downstream debiased machine learning estimators. We formalize random seed stability via a concentration condition and prove that subbagging…

Methodology · Statistics 2026-04-21 Nicholas Williams , Alejandro Schuler

The theory of Local Intrinsic Dimensionality (LID) has become a valuable tool for characterizing local complexity within and across data manifolds, supporting a range of data mining and machine learning tasks. Accurate LID estimation…

Machine Learning · Computer Science 2026-03-26 Kristóf Péter , Ricardo J. G. B. Campello , James Bailey , Michael E. Houle

Standard Bayesian inference is known to be sensitive to model misspecification, leading to unreliable uncertainty quantification and poor predictive performance. However, finding generally applicable and computationally feasible methods for…

Methodology · Statistics 2020-07-31 Jonathan H. Huggins , Jeffrey W. Miller

Series and polynomial regression are able to approximate the same function classes as neural networks. However, these methods are rarely used in practice, although they offer more interpretability than neural networks. In this paper, we…

Machine Learning · Statistics 2024-09-19 Sylvia Klosin , Jaume Vives-i-Bastida

We study the consistency of sample mean-variance portfolios of arbitrarily high dimension that are based on Bayesian or shrinkage estimation of the input parameters as well as weighted sampling. In an asymptotic setting where the number of…

Portfolio Management · Quantitative Finance 2015-05-30 Francisco Rubio , Xavier Mestre , Daniel P. Palomar

Conformal prediction is a generic methodology for finite-sample valid distribution-free prediction. This technique has garnered a lot of attention in the literature partly because it can be applied with any machine learning algorithm that…

Methodology · Statistics 2024-04-12 Yachong Yang , Arun Kumar Kuchibhotla

We provide a unified analysis of the predictive risk of ridge regression and regularized discriminant analysis in a dense random effects model. We work in a high-dimensional asymptotic regime where $p, n \to \infty$ and $p/n \to \gamma \in…

Statistics Theory · Mathematics 2015-11-05 Edgar Dobriban , Stefan Wager

Weighted empirical risk minimization is a common approach to prediction under distribution drift. This article studies its out-of-sample prediction error under nonstationarity. We provide a general decomposition of the excess risk into a…

Machine Learning · Statistics 2026-05-19 Tobias Brock , Thomas Nagler

We analyze the prediction error of ridge regression in an asymptotic regime where the sample size and dimension go to infinity at a proportional rate. In particular, we consider the role played by the structure of the true regression…

Statistics Theory · Mathematics 2021-03-09 Dominic Richards , Jaouad Mourtada , Lorenzo Rosasco