中文
相关论文

相关论文: Gradient descent in higher codimension

200 篇论文

When training neural networks, it has been widely observed that a large step size is essential in stochastic gradient descent (SGD) for obtaining superior models. However, the effect of large step sizes on the success of SGD is not well…

机器学习 · 计算机科学 2023-02-17 Amirkeivan Mohtashami , Martin Jaggi , Sebastian Stich

Many problems in high-dimensional statistics and optimization involve minimization over nonconvex constraints-for instance, a rank constraint for a matrix estimation problem-but little is known about the theoretical properties of such…

最优化与控制 · 数学 2017-10-20 Rina Foygel Barber , Wooseok Ha

For optimizing a non-convex function in finite dimension, a method is to add Brownian noise to a gradient descent, allowing for transitions between basins of attractions of different minimizers. To adapt this for optimization over a space…

概率论 · 数学 2025-05-13 Pierre Germain , Pierre Monmarché

A recent line of work has shown remarkable behaviors of the generalization error curves in simple learning models. Even the least-squares regression has shown atypical features such as the model-wise double descent, and further works have…

机器学习 · 统计学 2022-12-20 Antoine Bodin , Nicolas Macris

Distances between data points are widely used in machine learning applications. Yet, when corrupted by noise, these distances -- and thus the models based upon them -- may lose their usefulness in high dimensions. Indeed, the small marginal…

机器学习 · 计算机科学 2022-03-08 Robin Vandaele , Bo Kang , Tijl De Bie , Yvan Saeys

It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a…

机器学习 · 计算机科学 2020-06-29 Samuel L. Smith , Erich Elsen , Soham De

This paper is centered around the approximation of dynamical systems by means of Gaussian processes. To this end, trajectories of such systems must be collected to be used as training data. The measurements of these trajectories are…

系统与控制 · 电气工程与系统科学 2025-04-02 Tobias M. Wolff , Victor G. Lopez , Matthias A. Müller

Symmetries are prevalent in deep learning and can significantly influence the learning dynamics of neural networks. In this paper, we examine how exponential symmetries -- a broad subclass of continuous symmetries present in the model…

机器学习 · 计算机科学 2024-11-08 Liu Ziyin , Mingze Wang , Hongchao Li , Lei Wu

Numerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely unknown. This paper studies this fundamental problem through…

机器学习 · 计算机科学 2021-02-25 Tianyi Liu , Yan Li , Song Wei , Enlu Zhou , Tuo Zhao

Deep learning, a multi-layered neural network approach inspired by the brain, has revolutionized machine learning. One of its key enablers has been backpropagation, an algorithm that computes the gradient of a loss function with respect to…

Stochastic gradients for deep neural networks exhibit strong correlations along the optimization trajectory, and are often aligned with a small set of Hessian eigenvectors associated with outlier eigenvalues. Recent work shows that…

机器学习 · 计算机科学 2026-02-04 Julien Nicolas , Mohamed Maouche , Sonia Ben Mokhtar , Mark Coates

In this paper we give a description of the asymptotic behavior, as $\epsilon\to 0$, of the $\epsilon$-gradient flow in the finite dimensional case. Under very general assumptions we prove that it converges to an evolution obtained by…

泛函分析 · 数学 2007-05-23 Chiara Zanini

In this paper, we provide a theoretical study of noise geometry for minibatch stochastic gradient descent (SGD), a phenomenon where noise aligns favorably with the geometry of local landscape. We propose two metrics, derived from analyzing…

机器学习 · 计算机科学 2024-02-02 Mingze Wang , Lei Wu

In deep learning, it is common to use more network parameters than training points. In such scenarioof over-parameterization, there are usually multiple networks that achieve zero training error so that thetraining algorithm induces an…

机器学习 · 计算机科学 2023-08-22 Hung-Hsu Chou , Carsten Gieshoff , Johannes Maly , Holger Rauhut

The behavior of particles driven through a narrow constriction is investigated in experiment and simulation. The system of particles adapts to the confining potentials and the interaction energies by a self-consistent arrangement of the…

软凝聚态物质 · 物理学 2008-10-15 P. Henseler , A. Erbe , M. Köppl , P. Leiderer , P. Nielaba

We prove that stochastic gradient descent efficiently converges to the global optimizer of the maximum likelihood objective of an unknown linear time-invariant dynamical system from a sequence of noisy observations generated by the system.…

机器学习 · 计算机科学 2019-02-12 Moritz Hardt , Tengyu Ma , Benjamin Recht

Various results for higher-order perturbative calculations in the gradient-flow formalism are reviewed, including the gradient-flow beta function and the small-flow-time expansion of the hadronic vacuum polarization and the energy-momentum…

高能物理 - 格点 · 物理学 2024-11-21 Robert Harlander

Gauss-Newton methods and their stochastic version have been widely used in machine learning and signal processing. Their nonsmooth counterparts, modified Gauss-Newton or prox-linear algorithms, can lead to contrasting outcomes when compared…

最优化与控制 · 数学 2023-05-19 Krishna Pillutla , Vincent Roulet , Sham Kakade , Zaid Harchaoui

How to find flat minima? We propose running normalized gradient descent, usually reserved for nonsmooth optimization, with sufficiently slowly diminishing step sizes. This induces implicit regularization towards flat minima if an…

最优化与控制 · 数学 2026-02-10 Cédric Josz

Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averaging over large batches. Long plateaus in which one observes…

机器学习 · 计算机科学 2025-03-25 Raphaël Berthier , Andrea Montanari , Kangjie Zhou