English
Related papers

Related papers: Heavy-Ball Momentum Method in Continuous Time and …

200 papers

The stochastic heavy ball method (SHB), also known as stochastic gradient descent (SGD) with Polyak's momentum, is widely used in training neural networks. However, despite the remarkable success of such algorithm in practice, its…

Machine Learning · Computer Science 2023-02-07 Diyuan Wu , Vyacheslav Kungurtsev , Marco Mondelli

Motivated by the conspicuous use of momentum-based algorithms in deep learning, we study a nonsmooth nonconvex stochastic heavy ball method and show its convergence. Our approach builds upon semialgebraic (definable) assumptions commonly…

Optimization and Control · Mathematics 2024-01-24 Tam Le

The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-scale data. In this context, momentum-based stochastic optimisation algorithms have become…

Optimization and Control · Mathematics 2024-11-06 Kexin Jin , Jonas Latz , Chenguang Liu , Alessandro Scagliotti

The Heavy Ball Method, proposed by Polyak over five decades ago, is a first-order method for optimizing continuous functions. While its stochastic counterpart has proven extremely popular in training deep networks, there are almost no known…

Machine Learning · Computer Science 2021-02-16 Jun-Kun Wang , Jacob Abernethy

We propose a family of optimization methods that achieve linear convergence using first-order gradient information and constant step sizes on a class of convex functions much larger than the smooth and strongly convex ones. This larger…

Optimization and Control · Mathematics 2018-09-14 Chris J. Maddison , Daniel Paulin , Yee Whye Teh , Brendan O'Donoghue , Arnaud Doucet

Explicit, momentum-based dynamics that optimize functions defined on Lie groups can be constructed via variational optimization and momentum trivialization. Structure preserving time discretizations can then turn this dynamics into…

Machine Learning · Computer Science 2024-06-03 Lingkai Kong , Molei Tao

Combinatorial optimization (CO) has been a hot research topic because of its theoretic and practical importance. As a classic CO problem, deep hashing aims to find an optimal code for each data from finite discrete possibilities, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-14 Chaoyou Fu , Guoli Wang , Xiang Wu , Qian Zhang , Ran He

The heavy-ball momentum method accelerates gradient descent with a momentum term but lacks accelerated convergence for general smooth strongly convex problems. This work introduces the Accelerated Over-Relaxation Heavy-Ball (AOR-HB) method,…

Optimization and Control · Mathematics 2025-02-18 Jingrong Wei , Long Chen

We study distributed optimization to minimize a global objective that is a sum of smooth and strongly-convex local cost functions. Recently, several algorithms over undirected and directed graphs have been proposed that use a gradient…

Optimization and Control · Mathematics 2018-08-13 Ran Xin , Usman A. Khan

We propose heavy ball neural ordinary differential equations (HBNODEs), leveraging the continuous limit of the classical momentum accelerated gradient descent, to improve neural ODEs (NODEs) training and inference. HBNODEs have two…

Machine Learning · Computer Science 2021-10-12 Hedi Xia , Vai Suliafu , Hangjie Ji , Tan M. Nguyen , Andrea L. Bertozzi , Stanley J. Osher , Bao Wang

We reconsider the variational integration of optimal control problems for mechanical systems based on a direct discretization of the Lagrange-d'Alembert principle. This approach yields discrete dynamical constraints which by construction…

Optimization and Control · Mathematics 2012-04-30 C. M. Campos , O. Junge , S. Ober-Blöbaum

We present a simple and easy to implement method for the numerical solution of a rather general class of Hamilton-Jacobi-Bellman (HJB) equations. In many cases, the considered problems have only a viscosity solution, to which, fortunately,…

Computational Finance · Quantitative Finance 2011-02-17 Jan Hendrik Witte , Christoph Reisinger

Arguably, the two most popular accelerated or momentum-based optimization methods in machine learning are Nesterov's accelerated gradient and Polyaks's heavy ball, both corresponding to different discretizations of a particular second order…

Optimization and Control · Mathematics 2020-12-25 Guilherme França , Jeremias Sulam , Daniel P. Robinson , René Vidal

Heavy ball momentum is crucial in accelerating (stochastic) gradient-based optimization algorithms for machine learning. Existing heavy ball momentum is usually weighted by a uniform hyperparameter, which relies on excessive tuning.…

Machine Learning · Computer Science 2021-10-19 Tao Sun , Huaming Ling , Zuoqiang Shi , Dongsheng Li , Bao Wang

Optimization is an important module of modern machine learning applications. Tremendous efforts have been made to accelerate optimization algorithms. A common formulation is achieving a lower loss at a given time. This enables a…

Machine Learning · Computer Science 2025-05-29 Zhonglin Xie , Yiman Fong , Haoran Yuan , Zaiwen Wen

We propose a deep learning algorithm for high dimensional optimal stopping problems. Our method is inspired by the penalty method for solving free boundary PDEs. Within our approach, the penalized PDE is approximated using the Deep BSDE…

Mathematical Finance · Quantitative Finance 2026-04-07 Yunfei Peng , Pengyu Wei , Wei Wei

This paper is concerned with high moment and pathwise error estimates for both velocity and pressure approximations of the Euler-Maruyama scheme for time discretization and its two fully discrete mixed finite element discretizations. The…

Numerical Analysis · Mathematics 2021-07-01 Liet Vo

In this paper, we propose a unified framework, the Hessian discretisation method (HDM), which is based on four discrete elements (called altogether a Hessian discretisation) and a few intrinsic indicators of accuracy, independent of the…

Numerical Analysis · Mathematics 2018-08-28 Jérôme Droniou , Bishnu P. Lamichhane , Devika Shylaja

Despite the remarkable success of diffusion models in image generation, slow sampling remains a persistent issue. To accelerate the sampling process, prior studies have reformulated diffusion sampling as an ODE/SDE and introduced…

Computer Vision and Pattern Recognition · Computer Science 2023-07-24 Suttisak Wizadwongsa , Worameth Chinchuthakun , Pramook Khungurn , Amit Raj , Supasorn Suwajanakorn

Traditionally, the delay margin of a looped system is computed by considering both the controller and system representations that evolve in the same space (e.g. either continuous or discrete-time). However, as in practice the system is…

Systems and Control · Computer Science 2018-11-30 V. Bellet , C. Poussot-Vassal , C. Pagetti , T. Loquen