English
Related papers

Related papers: Sharp Risk Bounds for Early-Stopping in Gaussian L…

200 papers

We study the convergence dynamics of Gradient Descent (GD) in a minimal binary classification setting, consisting of a two-neuron ReLU network and two training instances. We prove that even under these strong simplifying assumptions, while…

Machine Learning · Computer Science 2026-03-03 Guy Smorodinsky , Sveta Gimpleson , Itay Safran

The support vector machine (SVM) and minimum Euclidean norm least squares regression are two fundamentally different approaches to fitting linear models, but they have recently been connected in models for very high-dimensional data through…

Machine Learning · Computer Science 2021-10-28 Navid Ardeshir , Clayton Sanford , Daniel Hsu

Error bound conditions (EBC) are properties that characterize the growth of an objective function when a point is moved away from the optimal set. They have recently received increasing attention in the field of optimization for developing…

Machine Learning · Statistics 2018-05-15 Mingrui Liu , Xiaoxuan Zhang , Lijun Zhang , Rong Jin , Tianbao Yang

The Polyak-Lojasiewicz (PL) inequality is a sufficient condition for establishing linear convergence of gradient descent, even in non-convex settings. While several recent works use a PL-based analysis to establish linear convergence of…

Machine Learning · Computer Science 2021-10-07 Adityanarayanan Radhakrishnan , Mikhail Belkin , Caroline Uhler

We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the ``near initialization'' perspective, we shed considerable new light on the…

Machine Learning · Computer Science 2022-10-03 Arindam Banerjee , Pedro Cisneros-Velarde , Libin Zhu , Mikhail Belkin

Early stopping of iterative algorithms is an algorithmic regularization method to avoid over-fitting in estimation and classification. In this paper, we show that early stopping can also be applied to obtain the minimax optimal testing in a…

Statistics Theory · Mathematics 2018-09-18 Meimei Liu , Guang Cheng

We consider the problem of nonparametric regression under shape constraints. The main examples include isotonic regression (with respect to any partial order), unimodal/convex regression, additive shape-restricted regression, and…

Statistics Theory · Mathematics 2018-07-03 Adityanand Guntuboyina , Bodhisattva Sen

In the first part of the paper, we study reflected backward stochastic differential equations (RBSDEs) with lower obstacle which is assumed to be right upper-semicontinuous but not necessarily right-continuous. We prove existence and…

Probability · Mathematics 2017-05-11 Miryana Grigorova , Peter Imkeller , Elias Offen , Youssef Ouknine , Marie-Claire Quenez

Saddle point problems, ubiquitous in optimization, extend beyond game theory to diverse domains like power networks and reinforcement learning. This paper presents novel approaches to tackle saddle point problem, with a focus on…

Optimization and Control · Mathematics 2024-04-09 Anik Kumar Paul , Arun D Mahindrakar , Rachel K Kalaimani

This work examines under what circumstances adaptivity for truncated SVD estimation can be achieved by an early stopping rule based on the smoothed residuals $ \| ( A A^{\top} )^{\alpha / 2} ( Y - A \hat{\mu}^{( m )}) \|^{2} $. Lower and…

Statistics Theory · Mathematics 2020-09-01 Bernhard Stankewitz

Stochastic gradient descent is one of the most successful approaches for solving large-scale problems, especially in machine learning and statistics. At each iteration, it employs an unbiased estimator of the full gradient computed from one…

Numerical Analysis · Mathematics 2018-12-05 Bangti Jin , Xiliang Lu

We introduce an alternative closed form lower bound on the Gaussian process ($\mathcal{GP}$) likelihood based on the R\'enyi $\alpha$-divergence. This new lower bound can be viewed as a convex combination of the Nystr\"om approximation and…

Machine Learning · Statistics 2023-07-04 Xubo Yue , Raed Kontar

Prior work (Klochkov $\&$ Zhivotovskiy, 2021) establishes at most $O\left(\log (n)/n\right)$ excess risk bounds via algorithmic stability for strongly-convex learners with high probability. We show that under the similar common assumptions…

Machine Learning · Computer Science 2025-10-31 Bowei Zhu , Shaojie Li , Mingyang Yi , Yong Liu

We study the behavior of the posterior distribution in high-dimensional Bayesian Gaussian linear regression models having $p\gg n$, with $p$ the number of predictors and $n$ the sample size. Our focus is on obtaining quantitative finite…

Statistics Theory · Mathematics 2014-01-06 Nate Strawn , Artin Armagan , Rayan Saab , Lawrence Carin , David Dunson

The Convex Gaussian Min-Max Theorem (CGMT) has emerged as a prominent theoretical tool for analyzing the precise stochastic behavior of various statistical estimators in the so-called high dimensional proportional regime, where the sample…

Statistics Theory · Mathematics 2022-06-28 Qiyang Han , Yandi Shen

Existing error-bound-based analyses for stochastic algorithms that exhibit certain descent properties, such as randomized coordinate descent and randomized projection methods, are often limited in scope and typically lead to overly…

Optimization and Control · Mathematics 2026-03-19 Zhichun Yang , Li Jiang , Tianxiang Liu , Man-Chung Yue

We analyze the statistical properties of generalized cross-validation (GCV) and leave-one-out cross-validation (LOOCV) applied to early-stopped gradient descent (GD) in high-dimensional least squares regression. We prove that GCV is…

Statistics Theory · Mathematics 2024-02-27 Pratik Patil , Yuchen Wu , Ryan J. Tibshirani

We present upper and lower bounds for the prediction error of the Lasso. For the case of random Gaussian design, we show that under mild conditions the prediction error of the Lasso is up to smaller order terms dominated by the prediction…

Statistics Theory · Mathematics 2018-04-04 Sara van de Geer

This paper proposes a theory for $\ell_1$-norm penalized high-dimensional $M$-estimators, with nonconvex risk and unrestricted domain. Under high-level conditions, the estimators are shown to attain the rate of convergence…

Statistics Theory · Mathematics 2022-04-14 Jad Beyhum , François Portier

In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event…

Machine Learning · Computer Science 2024-11-20 Steffen Dereich , Sebastian Kassing