English
Related papers

Related papers: Towards Quantifying the Preconditioning Effect of …

200 papers

Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training of deep neural networks in artificial intelligence (AI)…

Optimization and Control · Mathematics 2025-11-11 Steffen Dereich , Thang Do , Arnulf Jentzen , Philippe von Wurstemberger

The dynamic behavior of RMSprop and Adam algorithms is studied through a combination of careful numerical experiments and theoretical explanations. Three types of qualitative features are observed in the training loss curve: fast initial…

Machine Learning · Computer Science 2021-10-01 Chao Ma , Lei Wu , Weinan E

Langevin dynamics (LD) is widely used for sampling from distributions and for optimization. In this work, we derive a closed-form expression for the expected loss of preconditioned LD near stationary points of the objective function. We use…

Machine Learning · Computer Science 2024-02-22 Amitay Bar , Rotem Mulayoff , Tomer Michaeli , Ronen Talmon

Adam has become one of the most favored optimizers in deep learning problems. Despite its success in practice, numerous mysteries persist regarding its theoretical understanding. In this paper, we study the implicit bias of Adam in linear…

Machine Learning · Statistics 2024-06-18 Chenyang Zhang , Difan Zou , Yuan Cao

We address the slow convergence and poor stability of quasi-newton sequential quadratic programming (SQP) methods that is observed when solving experimental design problems, in particular when they are large. Our findings suggest that this…

Optimization and Control · Mathematics 2011-08-09 M. S. Mommer , A. Sommer , J. P. Schlöder , H. G. Bock

Let $H\subset \R^{d+1}$ be a compact, convex, analytic hypersurface of finite type with a smooth measure $\sigma $ on $H$. Let $\kappa$ denote the Gaussian curvature on $H$. We consider the oscillatory integral $(\kappa^{1/2}…

Classical Analysis and ODEs · Mathematics 2025-06-16 Sanghyuk Lee , Sewook Oh

This paper considers the problem for finding the $(\delta,\epsilon)$-Goldstein stationary point of Lipschitz continuous objective, which is a rich function class to cover a great number of important applications. We construct a zeroth-order…

Quantum Physics · Physics 2024-10-22 Chengchang Liu , Chaowen Guan , Jianhao He , John C. S. Lui

In previous literature, backward error analysis was used to find ordinary differential equations (ODEs) approximating the gradient descent trajectory. It was found that finite step sizes implicitly regularize solutions because terms…

Machine Learning · Computer Science 2024-06-18 Matias D. Cattaneo , Jason M. Klusowski , Boris Shigida

This paper proposes a novel preconditioned implicit-explicit algorithm enhanced with the extrapolation technique for non-convex optimization problems. The algorithm employs a third-order Adams-Bashforth scheme for the nonlinear and explicit…

Optimization and Control · Mathematics 2025-09-19 Kelin Wu , Hongpeng Sun

Scalable training of large models (like BERT and GPT-3) requires careful optimization rooted in model design, architecture, and system capabilities. From a system standpoint, communication has become a major bottleneck, especially on…

Machine Learning · Computer Science 2021-07-01 Hanlin Tang , Shaoduo Gan , Ammar Ahmad Awan , Samyam Rajbhandari , Conglong Li , Xiangru Lian , Ji Liu , Ce Zhang , Yuxiong He

The long time effect of nonlinear perturbation to oscillatory linear systems can be characterized by the averaging method, and we consider first-order averaging for its simplest applicability to high-dimensional problems. Instead of the…

Classical Analysis and ODEs · Mathematics 2018-12-05 Molei Tao

Here I present a small update to the bias-correction term in the Adam optimizer that has the advantage of making smaller gradient updates in the first several steps of training. With the default bias-correction, Adam may actually make…

Machine Learning · Computer Science 2021-10-25 John St John

The ADAM optimizer is exceedingly popular in the deep learning community. Often it works very well, sometimes it doesn't. Why? We interpret ADAM as a combination of two aspects: for each weight, the update direction is determined by the…

Machine Learning · Computer Science 2020-12-15 Lukas Balles , Philipp Hennig

We present a derivative-based algorithm for nonlinearly constrained optimization problems that is tolerant of inaccuracies in the data. The algorithm solves a semi-smooth set of nonlinear equations that are equivalent to the first-order…

Optimization and Control · Mathematics 2017-09-21 Jason E. Hicken , Pengfei Meng , Alp Dener

We introduce AlphaGrad, a memory-efficient, conditionally stateless optimizer addressing the memory overhead and hyperparameter complexity of adaptive methods like Adam. AlphaGrad enforces scale invariance via tensor-wise L2 gradient…

Machine Learning · Computer Science 2025-04-24 Soham Sane

The performance of optimization methods is often tied to the spectrum of the objective Hessian. Yet, conventional assumptions, such as smoothness, do often not enable us to make finely-grained convergence statements -- particularly not for…

Optimization and Control · Mathematics 2024-02-08 Nikita Doikov , Sebastian U. Stich , Martin Jaggi

The Adam optimizer is currently presumably the most popular optimization method in deep learning. In this article we develop an ODE based method to study the Adam optimizer in a fast-slow scaling regime. For fixed momentum parameters and…

Optimization and Control · Mathematics 2025-11-07 Steffen Dereich , Arnulf Jentzen , Sebastian Kassing

Newton-type methods are typically analyzed under Lipschitz continuity of the Hessian, an assumption that can fail for objectives with higher-order or polynomial growth. We introduce a class of nonlinearly preconditioned Newton methods that…

Optimization and Control · Mathematics 2026-05-14 Alexander Bodard , Panagiotis Patrinos

In this paper, we present several new results on minimizing a nonsmooth and nonconvex function under a Lipschitz condition. Recent work shows that while the classical notion of Clarke stationarity is computationally intractable up to some…

Optimization and Control · Mathematics 2022-11-08 Michael I. Jordan , Tianyi Lin , Manolis Zampetakis

We demonstrate the possibility of (sub)exponential quantum speedup via a quantum algorithm that follows an adiabatic path of a gapped Hamiltonian with no sign problem. This strengthens the superpolynomial separation recently proved by…

Quantum Physics · Physics 2020-11-20 András Gilyén , Umesh Vazirani
‹ Prev 1 4 5 6 7 8 10 Next ›