English
Related papers

Related papers: Towards Quantifying the Preconditioning Effect of …

200 papers

Using the decay along the diagonal of the matrix representing the perturbation with respect to the Hermite basis, we prove a reducibility result in $L^2(\mathbb{R})$ for the one-dimensional quantum harmonic oscillator perturbed by time…

Dynamical Systems · Mathematics 2025-09-03 Emanuele Haus , Zhiqiang Wang

We propose Adam-mini, an optimizer that achieves on par or better performance than AdamW with 50% less memory footprint. Adam-mini reduces memory by cutting down the learning rate resources in Adam (i.e., $1/\sqrt{v}$). By investigating the…

Machine Learning · Computer Science 2025-02-25 Yushun Zhang , Congliang Chen , Ziniu Li , Tian Ding , Chenwei Wu , Diederik P. Kingma , Yinyu Ye , Zhi-Quan Luo , Ruoyu Sun

Preconditioning is essential in iterative methods for solving linear systems. It is also the implicit objective in updating approximations of Jacobians in optimization methods, e.g.,in quasi-Newton methods. Motivated by the latter, we study…

Numerical Analysis · Mathematics 2024-12-24 Woosuk L. Jung , David Torregrosa-Belén , Henry Wolkowicz

Is the standard weight decay in AdamW truly optimal? Although AdamW decouples weight decay from adaptive gradient scaling, a fundamental conflict remains: the Radial Tug-of-War. In deep learning, gradients tend to increase parameter norms…

Machine Learning · Computer Science 2026-02-06 Hao Chen , Jinghui Yuan , Hanmin Zhang

Gradient descent (GD) based optimization methods are these days the standard tools to train deep neural networks in artificial intelligence systems. In optimization procedures in deep learning the employed optimizer is often not the…

Optimization and Control · Mathematics 2025-09-24 Steffen Dereich , Robin Graeber , Arnulf Jentzen , Adrian Riekert

We introduce ADAHESSIAN, a second order stochastic optimization algorithm which dynamically incorporates the curvature of the loss function via ADAptive estimates of the HESSIAN. Second order algorithms are among the most powerful…

Machine Learning · Computer Science 2021-04-30 Zhewei Yao , Amir Gholami , Sheng Shen , Mustafa Mustafa , Kurt Keutzer , Michael W. Mahoney

Pre-conditioning is a well-known concept that can significantly improve the convergence of optimization algorithms. For noise-free problems, where good pre-conditioners are not known a priori, iterative linear algebra methods offer one way…

Machine Learning · Computer Science 2019-02-21 Filip de Roos , Philipp Hennig

We revisit the problem of robust linear regression under Gaussian covariates with an unknown covariance matrix of condition number $\kappa$. For this fundamental problem, significant gaps remain in our understanding of the trade-offs among…

Data Structures and Algorithms · Computer Science 2026-05-19 Deeksha Adil , Jarosław Błasiok , Hongjie Chen , Deepak Narayanan Sridharan

Adiabatic quantum computation is based on the adiabatic evolution of quantum systems. We analyse a particular class of qauntum adiabatic evolutions where either the initial or final Hamiltonian is a one-dimensional projector Hamiltonian on…

Quantum Physics · Physics 2015-05-13 Avatar Tulsi

In stochastic zeroth-order optimization, a problem of practical relevance is understanding how to fully exploit the local geometry of the underlying objective function. We consider a fundamental setting in which the objective function is…

Machine Learning · Computer Science 2023-12-27 Qian Yu , Yining Wang , Baihe Huang , Qi Lei , Jason D. Lee

We study the Landau-Zener Problem for a decaying two-level-system described by a non-hermitean Hamiltonian, depending analytically on time. Use of a super-adiabatic basis allows to calculate the non-adiabatic transition probability P in the…

Quantum Physics · Physics 2009-11-13 R. Schilling , Mark Vogelsberger , D. A. Garanin

Ever since Reddi et al. 2018 pointed out the divergence issue of Adam, many new variants have been designed to obtain convergence. However, vanilla Adam remains exceptionally popular and it works well in practice. Why is there a gap between…

Machine Learning · Computer Science 2023-01-16 Yushun Zhang , Congliang Chen , Naichen Shi , Ruoyu Sun , Zhi-Quan Luo

With limited high-quality data and growing compute, multi-epoch training is gaining back its importance across sub-areas of deep learning. Adam(W), versions of which are go-to optimizers for many tasks such as next token prediction, has two…

Machine Learning · Computer Science 2026-05-11 Matias D. Cattaneo , Boris Shigida

The paper presents the formulation, implementation, and evaluation of the ArcGD optimiser. The evaluation is conducted initially on a non-convex benchmark function and subsequently on a real-world ML dataset. The initial comparative study…

Machine Learning · Computer Science 2026-03-25 Nikhil Verma , Joonas Linnosmaa , Leonardo Espinosa-Leal , Napat Vajragupta

We discuss the problem of numerically backpropagating Hessians through ordinary differential equations (ODEs) in various contexts and elucidate how different approaches may be favourable in specific situations. We discuss both theoretical…

Optimization and Control · Mathematics 2023-01-20 Axel Ciceri , Thomas Fischbacher

We consider a convex minimization problem for which the objective is the sum of a homogeneous polynomial of degree four and a linear term. Such task arises as a subproblem in algorithms for quadratic inverse problems with a…

Optimization and Control · Mathematics 2024-04-24 Radu-Alexandru Dragomir , Yurii Nesterov

We introduce AdamS, a simple yet effective alternative to Adam for large language model (LLM) pretraining and post-training. By leveraging a novel denominator, i.e., the root of weighted sum of squares of the momentum and the current…

Machine Learning · Computer Science 2025-05-23 Huishuai Zhang , Bohan Wang , Luoxin Chen

Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient descent, Adam can converge to a different solution with a…

Machine Learning · Computer Science 2021-08-26 Difan Zou , Yuan Cao , Yuanzhi Li , Quanquan Gu

Layer-wise preconditioning methods are a family of memory-efficient optimization algorithms that introduce preconditioners per axis of each layer's weight tensors. These methods have seen a recent resurgence, demonstrating impressive…

Machine Learning · Computer Science 2025-02-05 Thomas T. Zhang , Behrad Moniri , Ansh Nagwekar , Faraz Rahman , Anton Xue , Hamed Hassani , Nikolai Matni

We propose a new first-order method for minimizing nonconvex functions with a Lipschitz continuous gradient and Hessian. The proposed method is an accelerated gradient descent with two restart mechanisms and finds a solution where the…

Optimization and Control · Mathematics 2024-06-19 Naoki Marumo , Akiko Takeda