English
Related papers

Related papers: Ridge interpolators in correlated factor regressio…

200 papers

We study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the…

Machine Learning · Statistics 2024-06-24 Matteo Vilucchio , Emanuele Troiani , Vittorio Erba , Florent Krzakala

Networked data, in which every training example involves two objects and may share some common objects with others, is used in many machine learning tasks such as learning to rank and link prediction. A challenge of learning from networked…

Machine Learning · Computer Science 2017-11-23 Yuanhong Wang , Yuyi Wang , Xingwu Liu , Juhua Pu

We study a generic class of \emph{random optimization problems} (rops) and their typical behavior. The foundational aspects of the random duality theory (RDT), associated with rops, were discussed in \cite{StojnicRegRndDlt10}, where it was…

Probability · Mathematics 2023-12-04 Mihailo Stojnic

We consider linear regression problems with a varying number of random projections, where we provably exhibit a double descent curve for a fixed prediction problem, with a high-dimensional analysis based on random matrix theory. We first…

Machine Learning · Computer Science 2023-03-15 Francis Bach

Off-the-shelf machine learning algorithms for prediction such as regularized logistic regression cannot exploit the information of time-varying features without previously using an aggregation procedure of such sequential data. However,…

Applications · Statistics 2019-09-26 C. Gary Mena , Arno De Caigny , Kristof Coussement , Koen W. De Bock , Stefan Lessmann

We study asymptotic minimax problems for estimating a $d$-dimensional regression parameter over spheres of growing dimension ($d\to \infty$). Assuming that the data follows a linear model with Gaussian predictors and errors, we show that…

Statistics Theory · Mathematics 2016-01-18 Lee H. Dicker

Many common estimators in machine learning and causal inference are linear smoothers, where the prediction is a weighted average of the training outcomes. Some estimators, such as ordinary least squares and kernel ridge regression, allow…

Machine Learning · Computer Science 2026-04-02 David Arbour , Harsh Parikh , Bijan Niknam , Elizabeth Stuart , Kara Rudolph , Avi Feller

Multiple randomization designs (MRDs) are a class of experimental designs used to handle interference in two-sided marketplaces. We investigate regression adjustment strategies for estimating total, spillover, and direct effects in MRDs. We…

Methodology · Statistics 2026-03-23 Timothy Sudijono , Lihua Lei , Lorenzo Masoero , Suhas Vijaykumar , Guido Imbens , James McQueen

Marginal association summary statistics have attracted great attention in statistical genetics, mainly because the primary results of most genome-wide association studies (GWAS) are produced by marginal screening. In this paper, we study…

Methodology · Statistics 2019-11-25 Bingxin Zhao , Hongtu Zhu

Gradient information is widely useful and available in applications, and is therefore natural to include in the training of neural networks. Yet little is known theoretically about the impact of Sobolev training -- regression with both…

Machine Learning · Statistics 2025-11-06 Katharine E Fisher , Matthew TC Li , Youssef Marzouk , Timo Schorlepp

We first study the generalization error of models that use a fixed feature representation (frozen intermediate layers) followed by a trainable readout layer. This setting encompasses a range of architectures, from deep random-feature models…

Statistics Theory · Mathematics 2025-11-10 Yessin Moakher , Malik Tiomoko , Cosme Louart , Zhenyu Liao

Multivariate regression is a widespread computational technique that may give meaningless results if the explanatory variables are too numerous or highly collinear. Tikhonov regularization, or ridge regression, is a popular approach to…

Biomolecules · Quantitative Biology 2015-12-29 Ugo Bastolla , Yves Dehouck

The solution to empirical risk minimization with $f$-divergence regularization (ERM-$f$DR) is extended to constrained optimization problems, establishing conditions for equivalence between the solution and constraints. A dual formulation of…

Machine Learning · Statistics 2025-02-21 Francisco Daunas , Iñaki Esnaola , Samir M. Perlaza , Gholamali Aminian

The semiparametric regression models have attracted increasing attention owing to their robustness compared to their parametric counterparts. This paper discusses the efficiency bound for functional response models (FRM), an emerging class…

Methodology · Statistics 2022-05-18 Jinyuan Liu , Tuo Lin , Tian Chen , Xinlian Zhang , Xin M. Tu

Kernel methods represent one of the most powerful tools in machine learning to tackle problems expressed in terms of function values and derivatives due to their capability to represent and model complex relations. While these methods show…

Statistics Theory · Mathematics 2015-11-06 Bharath K. Sriperumbudur , Zoltan Szabo

We introduce Invariant Risk Minimization (IRM), a learning paradigm to estimate invariant correlations across multiple training distributions. To achieve this goal, IRM learns a data representation such that the optimal classifier, on top…

Machine Learning · Statistics 2020-03-31 Martin Arjovsky , Léon Bottou , Ishaan Gulrajani , David Lopez-Paz

We study the ridge regression (L2 regularized least squares) problem and its dual, which is also a ridge regression problem. We observe that the optimality conditions describing the primal and dual optimal solutions can be formulated in…

Numerical Analysis · Mathematics 2018-01-22 Ademir Alves Riberio , Peter Richtárik

Random forest regression (RF) is an extremely popular tool for the analysis of high-dimensional data. Nonetheless, its benefits may be lessened in sparse settings due to weak predictors, and a pre-estimation dimension reduction (targeting)…

Fisher discriminant analysis (FDA) is a widely used method for classification and dimensionality reduction. When the number of predictor variables greatly exceeds the number of observations, one of the alternatives for conventional FDA is…

Machine Learning · Statistics 2018-11-30 Agniva Chowdhury , Jiasen Yang , Petros Drineas

The use of cumulative incidence functions for characterizing the risk of one type of event in the presence of others has become increasingly popular over the past decade. The problems of modeling, estimation and inference have been treated…

Methodology · Statistics 2020-11-16 Youngjoo Cho , Annette M. Molinaro , Chen Hu , Robert L. Strawderman