English
Related papers

Related papers: Stable Anderson Acceleration for Deep Learning

200 papers

We consider the application of the type-I Anderson acceleration to solving general non-smooth fixed-point problems. By interleaving with safe-guarding steps, and employing a Powell-type regularization and a re-start checking for strong…

Optimization and Control · Mathematics 2018-08-14 Junzi Zhang , Brendan O'Donoghue , Stephen Boyd

The alternating direction method of multipliers (ADMM) is a popular approach for solving optimization problems that are potentially non-smooth and with hard constraints. It has been applied to various computer graphics applications,…

Graphics · Computer Science 2019-09-04 Juyong Zhang , Yue Peng , Wenqing Ouyang , Bailin Deng

The alternating direction multiplier method (ADMM) is widely used in computer graphics for solving optimization problems that can be nonsmooth and nonconvex. It converges quickly to an approximate solution, but can take a long time to…

Optimization and Control · Mathematics 2020-06-29 Wenqing Ouyang , Yue Peng , Yuxin Yao , Juyong Zhang , Bailin Deng

The expectation-maximization (EM) algorithm is a well-known iterative method for computing maximum likelihood estimates from incomplete data. Despite its numerous advantages, a main drawback of the EM algorithm is its frequently observed…

Computation · Statistics 2018-08-14 Nicholas C. Henderson , Ravi Varadhan

In this paper, we propose an Anderson-accelerated stochastic extragradient algorithm for solving a class of stochastic variational inequalities, by incorporating Anderson acceleration into the stochastic extragradient method under a…

Optimization and Control · Mathematics 2026-05-27 Xin Qu , Wei Bian , Xiaojun Chen

We present the Anderson Accelerated Primal-Dual Hybrid Gradient (AA-PDHG), a fixed-point-based framework designed to overcome the slow convergence of the standard PDHG method for the solution of linear programming (LP) problems. We…

Optimization and Control · Mathematics 2025-08-12 Yingxin Zhou , Stefano Cipolla , Phan Tu Vuong

Fixed-point solvers are ubiquitous in nonlinear PDEs, yet their progress collapses whenever the Jacobian at the solution carries an eigenvalue arbitrarily close to one. We ask whether such stagnation can be removed without storing long…

Numerical Analysis · Mathematics 2026-01-06 Francesco Alemanno

We present a novel two-level sketching extension of the Alternating Anderson-Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multi-physics simulations governed by discretized partial differential…

Numerical Analysis · Mathematics 2026-05-20 Nicolás A. Barnafi , Massimiliano Lupo Pasini

This paper considers the numerical solution of generalized Sylvester matrix equations, which arise in many scientific and engineering applications but remain challenging to solve efficiently, particularly when the coefficient matrices are…

Numerical Analysis · Mathematics 2026-04-20 Hongjia Chen , Chun-Hua Zhang , Zhongming Teng , Lei Du

In this work, we propose a generalized alternating Anderson acceleration method, a periodic scheme composed of $t$ fixed-point iteration steps, interleaved with $s$ steps of Anderson acceleration with window size $m$, to solve linear and…

Numerical Analysis · Mathematics 2026-02-02 Yunhui He , Santolo Leveque

Anderson Acceleration (AA) has been widely used to solve nonlinear fixed-point problems due to its rapid convergence. This work focuses on a variant of AA in which multiple Picard iterations are performed between each AA step, referred to…

Numerical Analysis · Mathematics 2025-07-15 Xue Feng , M. Paul Laiu , Thomas Strohmer

Physics-guided deep learning is an important prevalent research topic in scientific machine learning, which has tremendous potential in various complex applications including science and engineering. In these applications, data is expensive…

Numerical Analysis · Mathematics 2024-11-11 Qingping Zhou , Guixian Xu , Zhexin Wen , Hongqiao Wang

Training of deep neural networks (DNNs) frequently involves optimizing several millions or even billions of parameters. Even with modern computing architectures, the computational expense of DNN training can inhibit, for instance, network…

Machine Learning · Computer Science 2020-06-26 Mauricio E. Tano , Gavin D. Portwood , Jean C. Ragusa

Acceleration of first order methods is mainly obtained via inertial techniques \`a la Nesterov, or via nonlinear extrapolation. The latter has known a recent surge of interest, with successful applications to gradient and proximal gradient…

Machine Learning · Statistics 2021-10-29 Quentin Bertrand , Mathurin Massias

Decoupled learning is a branch of model parallelism which parallelizes the training of a network by splitting it depth-wise into multiple modules. Techniques from decoupled learning usually lead to stale gradient effect because of their…

Machine Learning · Computer Science 2020-12-08 Huiping Zhuang , Zhiping Lin , Kar-Ann Toh

We study the asymptotic convergence of AA($m$), i.e., Anderson acceleration with window size $m$ for accelerating fixed-point methods $x_{k+1}=q(x_{k})$, $x_k \in R^n$. Convergence acceleration by AA($m$) has been widely observed but is not…

Optimization and Control · Mathematics 2022-05-04 Hans De Sterck , Yunhui He

In this paper, we propose a novel Anderson's acceleration method to solve nonlinear equations, which does \emph{not} require a restart strategy to achieve numerical stability. We propose the greedy and random versions of our algorithm.…

Optimization and Control · Mathematics 2024-03-26 Haishan Ye , Dachao Lin , Xiangyu Chang , Zhihua Zhang

This paper proposes an extra gradient Anderson-accelerated algorithm for solving pseudomonotone variational inequalities, which uses the extra gradient scheme with line search to guarantee the global convergence and Anderson acceleration to…

Optimization and Control · Mathematics 2026-05-27 Xin Qu , Wei Bian , Xiaojun Chen

Neural network training is inherently sequential where the layers finish the forward propagation in succession, followed by the calculation and back-propagation of gradients (based on a loss function) starting from the last layer. The…

Machine Learning · Computer Science 2023-12-01 Vahid Janfaza , Shantanu Mandal , Farabi Mahmud , Abdullah Muzahid

This paper develops an efficient and robust solution technique for the steady Boussinesq model of non-isothermal flow using Anderson acceleration applied to a Picard iteration. After analyzing the fixed point operator associated with the…

Numerical Analysis · Mathematics 2020-04-15 Sara Pollock , Leo G. Rebholz , Mengying Xiao