English
Related papers

Related papers: A Dynamical Systems Perspective on Nesterov Accele…

200 papers

Even for the gradient descent (GD) method applied to neural network training, understanding its optimization dynamics, including convergence rate, iterate trajectories, function value oscillations, and especially its implicit acceleration,…

Machine Learning · Computer Science 2026-05-22 Alexander Tyurin

We consider model order reduction of parameterized Hamiltonian systems describing nondissipative phenomena, like wave-type and transport dominated problems. The development of reduced basis methods for such models is challenged by two main…

Numerical Analysis · Mathematics 2021-05-27 Cecilia Pagliantini

In this paper we classify the pathwise asymptotic behaviour of the discretisation of a general autonomous scalar differential equation which has a unique and globally stable equilibrium. The underlying continuous equation is subjected to a…

Probability · Mathematics 2013-10-10 John A. D. Appleby , Jian Cheng , Alexandra Rodkina

Alternating minimization (AM) procedures are practically efficient in many applications for solving convex and non-convex optimization problems. On the other hand, Nesterov's accelerated gradient is theoretically optimal first-order method…

Optimization and Control · Mathematics 2021-09-16 Sergey Guminov , Pavel Dvurechensky , Nazarii Tupitsa , Alexander Gasnikov

The paper addresses an error analysis of an Eulerian finite element method used for solving a linearized Navier--Stokes problem in a time-dependent domain. In this study, the domain's evolution is assumed to be known and independent of the…

Numerical Analysis · Mathematics 2024-08-26 Michael Neilan , Maxim Olshanskii

Accelerated gradient descent iterations are widely used in optimization. It is known that, in the continuous-time limit, these iterations converge to a second-order differential equation which we refer to as the accelerated gradient flow.…

Optimization and Control · Mathematics 2020-06-16 Mohammad Farazmand

We take a Hamiltonian-based perspective to generalize Nesterov's accelerated gradient descent and Polyak's heavy ball method to a broad class of momentum methods in the setting of (possibly) constrained minimization in Euclidean and…

Optimization and Control · Mathematics 2020-11-17 Jelena Diakonikolas , Michael I. Jordan

Computational multi-scale methods capitalize on a large time-scale separation to efficiently simulate slow dynamics over long time intervals. For stochastic systems, one often aims at resolving the statistics of the slowest dynamics. This…

Numerical Analysis · Mathematics 2021-05-14 Kristian Debrabant , Giovanni Samaey , Przemysław Zieliński

Acceleration and momentum are the de facto standard in modern applications of machine learning and optimization, yet the bulk of the work on implicit regularization focuses instead on unaccelerated methods. In this paper, we study the…

Machine Learning · Statistics 2022-01-21 Yue Sheng , Alnur Ali

Gradient descent and stochastic gradient descent are central to modern machine learning, yet their behavior under large step sizes remains theoretically unclear. Recent work suggests that acceleration often arises near the edge of…

Machine Learning · Computer Science 2026-03-02 Sacchit Kale , Piyushi Manupriya , Pierre Marion , Francis Bach , Anant Raj

This work proposes an Accelerated Primal-Dual Fixed-Point (APDFP) method that employs Nesterov type acceleration to solve composite problems of the form min f(x) + g(Bx), where g is nonsmooth and B is a linear operator. The APDFP features…

Optimization and Control · Mathematics 2025-11-04 Ya-Nan Zhu

In this thesis we develop a novel framework to study smooth and strongly convex optimization algorithms, both deterministic and stochastic. Focusing on quadratic functions we are able to examine optimization algorithms as a recursive…

Optimization and Control · Mathematics 2014-10-24 Yossi Arjevani

Nesterov SGD is widely used for training modern neural networks and other machine learning models. Yet, its advantages over SGD have not been theoretically clarified. Indeed, as we show in our paper, both theoretically and empirically,…

Machine Learning · Computer Science 2019-09-30 Chaoyue Liu , Mikhail Belkin

In this work, we approach the minimization of a continuously differentiable convex function under linear equality constraints by a second-order dynamical system with asymptotically vanishing damping term. The system is formulated in terms…

Optimization and Control · Mathematics 2021-06-24 Radu Ioan Bot , Dang-Khoa Nguyen

Nesterov's Accelerated Gradient (NAG) for optimization has better performance than its continuous time limit (noiseless kinetic Langevin) when a finite step-size is employed \citep{shi2021understanding}. This work explores the sampling…

Machine Learning · Computer Science 2022-06-22 Ruilin Li , Hongyuan Zha , Molei Tao

This monograph covers some recent advances in a range of acceleration techniques frequently used in convex optimization. We first use quadratic optimization problems to introduce two key families of methods, namely momentum and nested…

Optimization and Control · Mathematics 2024-09-26 Alexandre d'Aspremont , Damien Scieur , Adrien Taylor

Deterministic approaches using iterative optimisation have been historically successful in diffeomorphic image registration (DiffIR). Although these approaches are highly accurate, they typically carry a significant computational burden.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Alexander Thorley , Xi Jia , Hyung Jin Chang , Boyang Liu , Karina Bunting , Victoria Stoll , Antonio de Marvao , Declan P. O'Regan , Georgios Gkoutos , Dipak Kotecha , Jinming Duan

In a real Hilbert space, we consider two classical problems: the global minimization of a smooth and convex function $f$ (i.e., a convex optimization problem) and finding the zeros of a monotone and continuous operator $V$ (i.e., a monotone…

Optimization and Control · Mathematics 2025-04-23 Hedy Attouch , Radu Ioan Bot , David Alexander Hulett , Dang-Khoa Nguyen

Acceleration of first order methods is mainly obtained via inertial techniques \`a la Nesterov, or via nonlinear extrapolation. The latter has known a recent surge of interest, with successful applications to gradient and proximal gradient…

Machine Learning · Statistics 2021-10-29 Quentin Bertrand , Mathurin Massias

Asynchronous optimization algorithms often require delay bounds to prove their convergence, though these bounds can be difficult to obtain in practice. Existing algorithms that do not require delay bounds often converge slowly. Therefore,…

Optimization and Control · Mathematics 2025-08-12 Ellie Pond , Yichen Zhao , Matthew Hale