English
Related papers

Related papers: Quantitative Convergences of Lie Group Momentum Op…

200 papers

Due to its simplicity and outstanding ability to generalize, stochastic gradient descent (SGD) is still the most widely used optimization method despite its slow convergence. Meanwhile, adaptive methods have attracted rising attention of…

Optimization and Control · Mathematics 2020-06-15 Xunpeng Huang , Runxin Xu , Hao Zhou , Zhe Wang , Zhengyang Liu , Lei Li

We introduce two new particle-based algorithms for learning latent variable models via marginal maximum likelihood estimation, including one which is entirely tuning-free. Our methods are based on the perspective of marginal maximum…

Machine Learning · Statistics 2024-03-04 Louis Sharrock , Daniel Dodd , Christopher Nemeth

We take a Hamiltonian-based perspective to generalize Nesterov's accelerated gradient descent and Polyak's heavy ball method to a broad class of momentum methods in the setting of (possibly) constrained minimization in Euclidean and…

Optimization and Control · Mathematics 2020-11-17 Jelena Diakonikolas , Michael I. Jordan

A variational formulation of accelerated optimization on normed spaces was recently introduced by considering a specific family of time-dependent Bregman Lagrangian and Hamiltonian systems whose corresponding trajectories converge to the…

Optimization and Control · Mathematics 2022-01-11 Valentin Duruisseaux , Melvin Leok

Adaptive optimizers can reduce to normalized steepest descent (NSD) when only adapting to the current gradient, suggesting a close connection between the two algorithmic families. A key distinction between their analyses, however, lies in…

Machine Learning · Computer Science 2025-11-26 Shuo Xie , Tianhao Wang , Beining Wu , Zhiyuan Li

We study nonlinearly preconditioned gradient methods for smooth nonconvex optimization problems, focusing on sigmoid preconditioners that inherently perform a form of gradient clipping akin to the widely used gradient clipping technique.…

Optimization and Control · Mathematics 2025-10-14 Konstantinos Oikonomidis , Jan Quan , Panagiotis Patrinos

Adversarial formulations such as generative adversarial networks (GANs) have rekindled interest in two-player min-max games. A central obstacle in the optimization of such games is the rotational dynamics that hinder their convergence. In…

Machine Learning · Computer Science 2023-06-22 Reyhane Askari Hemmat , Amartya Mitra , Guillaume Lajoie , Ioannis Mitliagkas

Parallel stochastic gradient methods are gaining prominence in solving large-scale machine learning problems that involve data distributed across multiple nodes. However, obtaining unbiased stochastic gradients, which have been the focus of…

Machine Learning · Computer Science 2025-01-14 Ali Beikmohammadi , Sarit Khirirat , Sindri Magnússon

A heavy top with a fixed point and a rigid body in an ideal fluid are important examples of Hamiltonian systems on a dual to the semidirect product Lie algebra $e(n)=so(n)\ltimes\mathbb R^n$. We give a Lagrangian derivation of the…

Exactly Solvable and Integrable Systems · Physics 2015-06-26 Yuri B. Suris

Decentralized stochastic optimization has emerged as a fundamental paradigm for large-scale machine learning. However, practical implementations often rely on biased gradient estimators arising from communication compression or inexact…

Optimization and Control · Mathematics 2026-04-10 Qing Xu , Yiwei Liao , Wenqi Fan , Xingxing You , Songyi Dian

In this work, we analyze two of the most fundamental algorithms in geodesically convex optimization: Riemannian gradient descent and (possibly inexact) Riemannian proximal point. We quantify their rates of convergence and produce different…

Optimization and Control · Mathematics 2024-03-18 David Martínez-Rubio , Christophe Roux , Sebastian Pokutta

Langevin Dynamics has been extensively employed in global non-convex optimization due to the concentration of its stationary distribution around the global minimum of the potential function at low temperatures. In this paper, we propose to…

Optimization and Control · Mathematics 2023-05-22 Ryo Fujino

A variety of widely used optimization methods like SignSGD and Muon can be interpreted as instances of steepest descent under different norm-induced geometries. In this work, we study the implicit bias of mini-batch stochastic steepest…

Machine Learning · Computer Science 2026-02-13 Jichu Li , Xuan Tang , Difan Zou

We study convergence of the trajectories of the Heavy Ball dynamical system, with constant damping coefficient, in the framework of convex and non-convex smooth optimization. By using the Polyak-{\L}ojasiewicz condition, we derive new…

Optimization and Control · Mathematics 2022-01-27 Vassilis Apidopoulos , Nicolò Ginatta , Silvia Villa

Structured reinforcement learning and stochastic optimization often involve parameters evolving on matrix Lie groups such as rotations and rigid-body transformations. We establish a representation-optimization dichotomy for…

Optimization and Control · Mathematics 2026-03-27 Sooraj KC , Vivek Mishra

The gradient descent (GD) method -- is a fundamental and likely the most popular optimization algorithm in machine learning (ML), with a history traced back to a paper in 1847 (Cauchy, 1847). It was studied under various assumptions,…

Optimization and Control · Mathematics 2025-02-20 Aleksandr Lobanov , Alexander Gasnikov , Eduard Gorbunov , Martin Takáč

Langevin diffusion processes and their discretizations are often used for sampling from a target density. The most convenient framework for assessing the quality of such a sampling scheme corresponds to smooth and strongly log-concave…

Probability · Mathematics 2018-12-27 Arnak S. Dalalyan , Lionel Riou-Durand

Recently a majorization method for optimizing partition functions of log-linear models was proposed alongside a novel quadratic variational upper-bound. In the batch setting, it outperformed state-of-the-art first- and second-order…

Machine Learning · Computer Science 2013-09-24 Anna Choromanska , Tony Jebara

Mini-batch algorithms have been proposed as a way to speed-up stochastic convex optimization problems. We study how such algorithms can be improved using accelerated gradient methods. We provide a novel analysis, which shows how standard…

Machine Learning · Computer Science 2011-06-24 Andrew Cotter , Ohad Shamir , Nathan Srebro , Karthik Sridharan

We present a numerical discretisation of the coupled moment systems, previously introduced in Dahm and Helzel, which approximate the kinetic multi-scale model by Helzel and Tzavaras for sedimentation in suspensions of rod-like particles for…

Numerical Analysis · Mathematics 2024-01-29 Sina Dahm , Jan Giesselmann , Christiane Helzel