English
Related papers

Related papers: Foolish Crowds Support Benign Overfitting

200 papers

We introduce a randomly extrapolated primal-dual coordinate descent method that adapts to sparsity of the data matrix and the favorable structures of the objective function. Our method updates only a subset of primal and dual variables with…

Optimization and Control · Mathematics 2020-07-14 Ahmet Alacaoglu , Olivier Fercoq , Volkan Cevher

A regression model with more parameters than data points in the training data is overparametrized and has the capability to interpolate the training data. Based on the classical bias-variance tradeoff expressions, it is commonly assumed…

Machine Learning · Computer Science 2023-04-18 Tomas McKelvey

This paper examines fundamental error characteristics for a general class of matrix completion problems, where the matrix of interest is a product of two a priori unknown matrices, one of which is sparse, and the observations are noisy. Our…

Information Theory · Computer Science 2017-10-27 Abhinav V. Sambasivan , Jarvis D. Haupt

This paper studies the binary classification of two distributions with the same Gaussian copula in high dimensions. Under this semiparametric Gaussian copula setting, we derive an accurate semiparametric estimator of the log density ratio,…

Statistics Theory · Mathematics 2014-11-12 Yue Zhao , Marten Wegkamp

Variational Bayesian posterior inference often requires simplifying approximations such as mean-field parametrisation to ensure tractability. However, prior work has associated the variational mean-field approximation for Bayesian neural…

Machine Learning · Computer Science 2022-10-07 Richard Kurle , Ralf Herbrich , Tim Januschowski , Yuyang Wang , Jan Gasthaus

Sparse linear regression with ill-conditioned Gaussian random designs is widely believed to exhibit a statistical/computational gap, but there is surprisingly little formal evidence for this belief, even in the form of examples that are…

Data Structures and Algorithms · Computer Science 2022-03-08 Jonathan A. Kelner , Frederic Koehler , Raghu Meka , Dhruv Rohatgi

Let A be an n by m matrix with m>n, and suppose that the underdetermined linear system As=x admits a sparse solution s0 for which ||s0||_0 < 1/2 spark(A). Such a sparse solution is unique due to a well-known uniqueness theorem. Suppose now…

Information Theory · Computer Science 2016-11-17 Massoud Babaie-Zadeh , Christian Jutten , Hosein Mohimani

This paper considers the penalized least squares estimator with arbitrary convex penalty. When the observation noise is Gaussian, we show that the prediction error is a subgaussian random variable concentrated around its median. We apply…

Statistics Theory · Mathematics 2016-09-22 Pierre C. Bellec , Alexandre B. Tsybakov

It has recently been shown that for compressive sensing, significantly fewer measurements may be required if the sparsity assumption is replaced by the assumption the unknown vector lies near the range of a suitably-chosen generative model.…

Information Theory · Computer Science 2020-03-11 Zhaoqiang Liu , Jonathan Scarlett

Generative models that maximize model likelihood have gained traction in many practical settings. Among them, perturbation based approaches underpin many strong likelihood estimation models, yet they often face slow convergence and limited…

Information Theory · Computer Science 2025-10-27 Yirong Shen , Lu Gan , Cong Ling

Consider a regression model with fixed design and Gaussian noise where the regression function can potentially be well approximated by a function that admits a sparse representation in a given dictionary. This paper resorts to exponential…

Statistics Theory · Mathematics 2013-01-08 Philippe Rigollet , Alexandre B. Tsybakov

The phenomenon of benign overfitting is one of the key mysteries uncovered by deep learning methodology: deep neural networks seem to predict well, even with a perfect fit to noisy training data. Motivated by this phenomenon, we consider…

Machine Learning · Statistics 2022-06-08 Peter L. Bartlett , Philip M. Long , Gábor Lugosi , Alexander Tsigler

We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense…

Machine Learning · Statistics 2026-05-25 John Sous , Michael Winer

The problem of sparse linear regression is relevant in the context of linear system identification from large datasets. When data are collected from real-world experiments, measurements are always affected by perturbations or low-precision…

Optimization and Control · Mathematics 2020-04-01 S. M. Fosson , V. Cerone , D. Regruto

Overfitting in linear regression is broken down into two main causes. First, the formula for the estimator includes 'forbidden knowledge' about training observations' residuals, and it loses this advantage when deployed out-of-sample.…

Methodology · Statistics 2022-09-27 Chris Rohlfs

A number of open problems hinder our present ability to extract scientific information from data that will be gathered by the near-future gravitational-wave mission LISA. Many of these relate to the modeling, detection and characterization…

Instrumentation and Methods for Astrophysics · Physics 2020-02-17 Alvin J. K. Chua , Natalia Korsakova , Christopher J. Moore , Jonathan R. Gair , Stanislav Babak

This paper studies offline reinforcement learning with linear function approximation in a setting with decision-theoretic, but not estimation sparsity. The structural restrictions of the data-generating process presume that the transitions…

Machine Learning · Statistics 2024-01-24 Angela Zhou

We study the problem of designing minimax procedures in linear regression under the quantile risk. We start by considering the realizable setting with independent Gaussian noise, where for any given noise level and distribution of inputs,…

Statistics Theory · Mathematics 2024-06-19 Ayoub El Hanchi , Chris J. Maddison , Murat A. Erdogdu

We study the asymptotic generalization of an overparameterized linear model for multiclass classification under the Gaussian covariates bi-level model introduced in Subramanian et al.~'22, where the number of data points, features, and…

Machine Learning · Computer Science 2025-03-28 David X. Wu , Anant Sahai

Numerous recent works show that overparameterization implicitly reduces variance for min-norm interpolators and max-margin classifiers. These findings suggest that ridge regularization has vanishing benefits in high dimensions. We challenge…

Machine Learning · Statistics 2021-12-20 Konstantin Donhauser , Alexandru Ţifrea , Michael Aerni , Reinhard Heckel , Fanny Yang
‹ Prev 1 4 5 6 7 8 10 Next ›