English
Related papers

Related papers: Gradient descent in higher codimension

200 papers

When training neural networks, it has been widely observed that a large step size is essential in stochastic gradient descent (SGD) for obtaining superior models. However, the effect of large step sizes on the success of SGD is not well…

Machine Learning · Computer Science 2023-02-17 Amirkeivan Mohtashami , Martin Jaggi , Sebastian Stich

Many problems in high-dimensional statistics and optimization involve minimization over nonconvex constraints-for instance, a rank constraint for a matrix estimation problem-but little is known about the theoretical properties of such…

Optimization and Control · Mathematics 2017-10-20 Rina Foygel Barber , Wooseok Ha

For optimizing a non-convex function in finite dimension, a method is to add Brownian noise to a gradient descent, allowing for transitions between basins of attractions of different minimizers. To adapt this for optimization over a space…

Probability · Mathematics 2025-05-13 Pierre Germain , Pierre Monmarché

A recent line of work has shown remarkable behaviors of the generalization error curves in simple learning models. Even the least-squares regression has shown atypical features such as the model-wise double descent, and further works have…

Machine Learning · Statistics 2022-12-20 Antoine Bodin , Nicolas Macris

Distances between data points are widely used in machine learning applications. Yet, when corrupted by noise, these distances -- and thus the models based upon them -- may lose their usefulness in high dimensions. Indeed, the small marginal…

Machine Learning · Computer Science 2022-03-08 Robin Vandaele , Bo Kang , Tijl De Bie , Yvan Saeys

It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a…

Machine Learning · Computer Science 2020-06-29 Samuel L. Smith , Erich Elsen , Soham De

This paper is centered around the approximation of dynamical systems by means of Gaussian processes. To this end, trajectories of such systems must be collected to be used as training data. The measurements of these trajectories are…

Systems and Control · Electrical Eng. & Systems 2025-04-02 Tobias M. Wolff , Victor G. Lopez , Matthias A. Müller

Symmetries are prevalent in deep learning and can significantly influence the learning dynamics of neural networks. In this paper, we examine how exponential symmetries -- a broad subclass of continuous symmetries present in the model…

Machine Learning · Computer Science 2024-11-08 Liu Ziyin , Mingze Wang , Hongchao Li , Lei Wu

Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely unknown. This paper studies this fundamental problem through…

Machine Learning · Computer Science 2021-02-25 Tianyi Liu , Yan Li , Song Wei , Enlu Zhou , Tuo Zhao

Deep learning, a multi-layered neural network approach inspired by the brain, has revolutionized machine learning. One of its key enablers has been backpropagation, an algorithm that computes the gradient of a loss function with respect to…

Stochastic gradients for deep neural networks exhibit strong correlations along the optimization trajectory, and are often aligned with a small set of Hessian eigenvectors associated with outlier eigenvalues. Recent work shows that…

Machine Learning · Computer Science 2026-02-04 Julien Nicolas , Mohamed Maouche , Sonia Ben Mokhtar , Mark Coates

In this paper we give a description of the asymptotic behavior, as $\epsilon\to 0$, of the $\epsilon$-gradient flow in the finite dimensional case. Under very general assumptions we prove that it converges to an evolution obtained by…

Functional Analysis · Mathematics 2007-05-23 Chiara Zanini

In this paper, we provide a theoretical study of noise geometry for minibatch stochastic gradient descent (SGD), a phenomenon where noise aligns favorably with the geometry of local landscape. We propose two metrics, derived from analyzing…

Machine Learning · Computer Science 2024-02-02 Mingze Wang , Lei Wu

In deep learning, it is common to use more network parameters than training points. In such scenarioof over-parameterization, there are usually multiple networks that achieve zero training error so that thetraining algorithm induces an…

Machine Learning · Computer Science 2023-08-22 Hung-Hsu Chou , Carsten Gieshoff , Johannes Maly , Holger Rauhut

The behavior of particles driven through a narrow constriction is investigated in experiment and simulation. The system of particles adapts to the confining potentials and the interaction energies by a self-consistent arrangement of the…

Soft Condensed Matter · Physics 2008-10-15 P. Henseler , A. Erbe , M. Köppl , P. Leiderer , P. Nielaba

We prove that stochastic gradient descent efficiently converges to the global optimizer of the maximum likelihood objective of an unknown linear time-invariant dynamical system from a sequence of noisy observations generated by the system.…

Machine Learning · Computer Science 2019-02-12 Moritz Hardt , Tengyu Ma , Benjamin Recht

Various results for higher-order perturbative calculations in the gradient-flow formalism are reviewed, including the gradient-flow beta function and the small-flow-time expansion of the hadronic vacuum polarization and the energy-momentum…

High Energy Physics - Lattice · Physics 2024-11-21 Robert Harlander

Gauss-Newton methods and their stochastic version have been widely used in machine learning and signal processing. Their nonsmooth counterparts, modified Gauss-Newton or prox-linear algorithms, can lead to contrasting outcomes when compared…

Optimization and Control · Mathematics 2023-05-19 Krishna Pillutla , Vincent Roulet , Sham Kakade , Zaid Harchaoui

How to find flat minima? We propose running normalized gradient descent, usually reserved for nonsmooth optimization, with sufficiently slowly diminishing step sizes. This induces implicit regularization towards flat minima if an…

Optimization and Control · Mathematics 2026-02-10 Cédric Josz

Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averaging over large batches. Long plateaus in which one observes…

Machine Learning · Computer Science 2025-03-25 Raphaël Berthier , Andrea Montanari , Kangjie Zhou