English
Related papers

Related papers: Dimension-Free Bounds for Generalized First-Order …

200 papers

We focus on nonconvex and nonsmooth minimization problems with a composite objective, where the differentiable part of the objective is freed from the usual and restrictive global Lipschitz gradient continuity assumption. This longstanding…

Optimization and Control · Mathematics 2017-06-21 Jérôme Bolte , Shoham Sabach , Marc Teboulle , Yakov Vaisbourd

Infinitely wide or deep neural networks (NNs) with independent and identically distributed (i.i.d.) parameters have been shown to be equivalent to Gaussian processes. Because of the favorable properties of Gaussian processes, this…

Machine Learning · Computer Science 2026-03-24 Steven Adams , Andrea Patanè , Morteza Lahijanian , Luca Laurenti

We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression provides a strong theoretical foundation, traditional…

Machine Learning · Computer Science 2026-03-10 Davide Maran , Marcello Restelli

Approximate message passing algorithm enjoyed considerable attention in the last decade. In this paper we introduce a variant of the AMP algorithm that takes into account glassy nature of the system under consideration. We coin this…

Disordered Systems and Neural Networks · Physics 2019-02-07 Fabrizio Antenucci , Florent Krzakala , Pierfrancesco Urbani , Lenka Zdeborová

Distributed algorithms for solving additive or consensus optimization problems commonly rely on first-order or proximal splitting methods. These algorithms generally come with restrictive assumptions and at best enjoy a linear convergence…

Optimization and Control · Mathematics 2017-05-11 Sina Khoshfetrat Pakazad , Christian A. Naesseth , Fredrik Lindsten , Anders Hansson

Generalized Linear Models (GLMs), where a random vector $\mathbf{x}$ is observed through a noisy, possibly nonlinear, function of a linear transform $\mathbf{z}=\mathbf{Ax}$ arise in a range of applications in nonlinear filtering and…

Information Theory · Computer Science 2016-05-03 Sundeep Rangan , Alyson K. Fletcher , Philip Schniter , Ulugbek Kamilov

We investigate the minimal error in approximating a general probability measure $\mu$ on $\mathbb{R}^d$ by the uniform measure on a finite set with prescribed cardinality $n$. The error is measured in the $p$-Wasserstein distance. In…

Probability · Mathematics 2024-08-26 Filippo Quattrocchi

This work discusses how to derive upper bounds for the expected generalisation error of supervised learning algorithms by means of the chaining technique. By developing a general theoretical framework, we establish a duality between…

Machine Learning · Statistics 2022-07-01 Eugenio Clerico , Amitis Shidani , George Deligiannidis , Arnaud Doucet

We consider the problem of Gaussian multiplier bootstrap procedures for the $k$th largest statistics and functions of the top $k$ order statistics, which are commonly encountered in high-dimensional statistical inference. Such a problem has…

Statistics Theory · Mathematics 2026-03-04 Yixi Ding , Qizhai Li , Yuke Shi , Liuquan Sun , Luobin Zhang

Gaussian processes (GPs) offer a principled probabilistic model over functions, but exact inference is restricted to the linear-Gaussian regime. We establish an explicit equivalence between GPs and a class of linear diffusion models,…

Gaussian processes (GPs) offer a flexible class of priors for nonparametric Bayesian regression, but popular GP posterior inference methods are typically prohibitively slow or lack desirable finite-data guarantees on quality. We develop an…

Machine Learning · Statistics 2019-03-28 Jonathan H. Huggins , Trevor Campbell , Mikołaj Kasprzak , Tamara Broderick

We derive first-order (in the stepsize) bounds on the bias in Wasserstein distances of the invariant measure of stochastic gradient kinetic Langevin dynamics with minimal assumptions on the stochastic gradient noise. These bounds sharpen…

Computation · Statistics 2026-04-28 Daniel Paulin , Peter A. Whalley

In this paper, we establish the non-asymptotic validity of the multiplier bootstrap procedure for constructing the confidence sets using the Stochastic Gradient Descent (SGD) algorithm. Under appropriate regularity conditions, our approach…

Deep neural networks generalize well despite being heavily overparameterized, in apparent contradiction with classical learning theory based on uniform convergence over fixed hypothesis spaces. Uniform bounds over the entire parameter space…

Machine Learning · Statistics 2026-05-15 Hubert Leroux , Jean Marcus , Julien Roger

We provide the detailed asymptotic behavior for first-order aggregation models of heterogeneous oscillators. Due to the dissimilarity of natural frequencies, one could expect that all relative distances converge to definite positive value…

Dynamical Systems · Mathematics 2022-06-03 Dohyun Kim , Hansol Park

We exploit analogies between first-order algorithms for constrained optimization and non-smooth dynamical systems to design a new class of accelerated first-order algorithms for constrained optimization. Unlike Frank-Wolfe or projected…

Optimization and Control · Mathematics 2025-05-02 Michael Muehlebach , Michael I. Jordan

Maxwell-Amp\`{e}re-Nernst-Planck (MANP) equations were recently proposed to model the dynamics of charged particles. In this study, we enhance a numerical algorithm of this system with deep learning tools. The proposed hybrid algorithm…

Numerical Analysis · Mathematics 2023-12-12 Cheng Chang , Zhouping Xin , Tieyong Zeng

The ability of overparameterized deep networks to generalize well has been linked to the fact that stochastic gradient descent (SGD) finds solutions that lie in flat, wide minima in the training loss -- minima where the output of the…

Machine Learning · Computer Science 2019-06-03 Vaishnavh Nagarajan , J. Zico Kolter

Many recent studies on first-order methods (FOMs) focus on \emph{composite non-convex non-smooth} optimization with linear and/or nonlinear function constraints. Upper (or worst-case) complexity bounds have been established for these…

Optimization and Control · Mathematics 2023-07-18 Wei Liu , Qihang Lin , Yangyang Xu

Dense associative memories (DAMs) store and retrieve patterns via energy-function based fixed points, but existing models are limited to vector representations. We extend DAMs to Gaussian densities equipped with the 2-Wasserstein distance.…

Machine Learning · Computer Science 2026-02-03 Chandan Tankala , Krishnakumar Balasubramanian