English
Related papers

Related papers: Precise Asymptotics of Bagging Regularized M-estim…

200 papers

A common approach to statistical learning with big-data is to randomly split it among $m$ machines and learn the parameter of interest by averaging the $m$ individual estimates. In this paper, focusing on empirical risk minimization, or…

Machine Learning · Statistics 2016-06-14 Jonathan Rosenblatt , Boaz Nadler

Regularized linear regression is central to machine learning, yet its high-dimensional behavior with informative priors remains poorly understood. We provide the first exact asymptotic characterization of training and test risks for maximum…

Machine Learning · Statistics 2026-01-28 Malik Tiomoko , Ekkehard Schnoor

Optimization under uncertainty and risk is indispensable in many practical situations. Our paper addresses stability of optimization problems using composite risk functionals which are subjected to measure perturbations. Our main focus is…

Optimization and Control · Mathematics 2022-01-06 Darinka Dentcheva , Yang Lin , Spiridon Penev

Stacking regressions is an ensemble technique that forms linear combinations of different regression estimators to enhance predictive accuracy. The conventional approach uses cross-validation data to generate predictions from the…

Machine Learning · Statistics 2024-10-10 Xin Chen , Jason M. Klusowski , Yan Shuo Tan

Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial…

Methodology · Statistics 2025-10-08 Amalan Mahendran , Helen Thompson , James M. McGree

Many regularization schemes for high-dimensional regression have been put forward. Most require the choice of a tuning parameter, using model selection criteria or cross-validation schemes. We show that a simple non-negative or…

Methodology · Statistics 2012-02-07 Nicolai Meinshausen

The purpose of this article is to develop a general parametric estimation theory that allows the derivation of the limit distribution of estimators in non-regular models where the true parameter value may lie on the boundary of the…

Statistics Theory · Mathematics 2022-11-28 Junichiro Yoshida , Nakahiro Yoshida

Analysis of non-asymptotic estimation error and structured statistical recovery based on norm regularized regression, such as Lasso, needs to consider four aspects: the norm, the loss function, the design matrix, and the noise model. This…

Machine Learning · Statistics 2015-12-01 Arindam Banerjee , Sheng Chen , Farideh Fazayeli , Vidyashankar Sivakumar

This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a…

Machine Learning · Statistics 2026-05-12 Lucas Morisset , Alain Durmus , Adrien Hardy

It is increasingly common to encounter prediction tasks in the biomedical sciences for which multiple datasets are available for model training. Common approaches such as pooling datasets and applying standard statistical learning methods…

Machine Learning · Statistics 2021-10-05 Gabriel Loewinger , Rolando Acosta Nunez , Rahul Mazumder , Giovanni Parmigiani

The maximum likelihood estimator (MLE) is pivotal in statistical inference, yet its application is often hindered by the absence of closed-form solutions for many models. This poses challenges in real-time computation scenarios,…

Methodology · Statistics 2025-04-16 Pedro L. Ramos , Eduardo Ramos , Francisco A. Rodrigues , Francisco Louzada

We compute precise asymptotic expressions for the learning curves of least squares random feature (RF) models with either a separable strongly convex regularization or the $\ell_1$ regularization. We propose a novel multi-level application…

Machine Learning · Statistics 2023-03-02 David Bosch , Ashkan Panahi , Ayca Özcelikkale , Devdatt Dubhash

Language models can learn a range of capabilities from unsupervised training on text corpora. However, to solve a particular problem (such as text summarization) it is typically necessary to fine-tune them on a task-specific dataset. It is…

Computation and Language · Computer Science 2022-03-16 Adam Gleave , Geoffrey Irving

We present a framework for the theoretical analysis of ensembles of low-complexity empirical risk minimisers trained on independent random compressions of high-dimensional data. First we introduce a general distribution-dependent…

Machine Learning · Computer Science 2021-06-03 Henry W. J. Reeve , Ata Kaban

Major progress has been made in the previous decade to characterize the asymptotic behavior of regularized M-estimators in high-dimensional regression problems in the proportional asymptotic regime where the sample size $n$ and the number…

Statistics Theory · Mathematics 2024-10-15 Pierre C. Bellec , Takuya Koriyama

In this paper we compare and contrast the behavior of the posterior predictive distribution to the risk of the maximum a posteriori estimator for the random features regression model in the overparameterized regime. We will focus on the…

Machine Learning · Statistics 2023-10-30 Youngsoo Baek , Samuel I. Berchuck , Sayan Mukherjee

Majorization-minimization (MM) is a standard iterative optimization technique which consists in minimizing a sequence of convex surrogate functionals. MM approaches have been particularly successful to tackle inverse problems and…

Applications · Statistics 2018-08-01 Yousra Bekhti , Felix Lucka , Joseph Salmon , Alexandre Gramfort

Sparse regression and classification estimators that respect group structures have application to an assortment of statistical and machine learning problems, from multitask learning to sparse additive modeling to hierarchical selection.…

Methodology · Statistics 2024-03-11 Ryan Thompson , Farshid Vahid

This paper explores the estimation of a panel data model with cross-sectional interaction that is flexible both in its approach to specifying the network of connections between cross-sectional units, and in controlling for unobserved…

Econometrics · Economics 2021-11-23 Ayden Higgins , Federico Martellosio

In this paper, a general class of regularized $M$-estimators of scatter matrix are proposed which are suitable also for low or insufficient sample support (small $n$ and large $p$) problems. The considered class constitutes a natural…

Applications · Statistics 2015-06-19 Esa Ollila , David E. Tyler