English
Related papers

Related papers: Stable Anderson Acceleration for Deep Learning

200 papers

This paper applies the Anderson Acceleration (AA) technique to accelerate the Fenchel dual gradient method (FDGM) to solve constrained optimization problems over time-varying networks. AA is originally designed for accelerating fixed-point…

Optimization and Control · Mathematics 2026-01-21 Haijuan Liu , Xuyang Wu

Anderson acceleration is an effective technique for enhancing the efficiency of fixed-point iterations; however, analyzing its convergence in nonsmooth settings presents significant challenges. In this paper, we investigate a class of…

Optimization and Control · Mathematics 2024-10-16 Kexin Li , Luwei Bai , Xiao Wang , Hao Wang

In this paper, we propose and analyze a set of fully non-stationary Anderson acceleration algorithms with dynamic window sizes and optimized damping. Although Anderson acceleration (AA) has been used for decades to speed up nonlinear…

Numerical Analysis · Mathematics 2022-03-29 Kewang Chen , Cornelis Vuik

Two adaptive relaxation strategies are proposed for Anderson acceleration. They are specifically designed for applications in which mappings converge to a fixed point. Their superiority over alternative Anderson acceleration is demonstrated…

Numerical Analysis · Mathematics 2024-09-02 Nicolas Lepage-Saucier

Stochastic gradient descent (SGD) and its many variants are the widespread optimization algorithms for training deep neural networks. However, SGD suffers from inevitable drawbacks, including vanishing gradients, lack of theoretical…

Machine Learning · Computer Science 2024-01-09 Zeinab Ebrahimi , Gustavo Batista , Mohammad Deghat

A pervasive approach in scientific computing is to express the solution to a given problem as the limit of a sequence of vectors or other mathematical objects. In many situations these sequences are generated by slowly converging iterative…

Numerical Analysis · Mathematics 2025-07-17 Yousef Saad

Anderson mixing (AM) is an acceleration method for fixed-point iterations. Despite its success and wide usage in scientific computing, the convergence theory of AM remains unclear, and its applications to machine learning problems are not…

Machine Learning · Computer Science 2021-10-05 Fuchao Wei , Chenglong Bao , Yang Liu

The derivative-free projection method (DFPM) is an efficient algorithm for solving monotone nonlinear equations. As problems grow larger, there is a strong demand for speeding up the convergence of DFPM. This paper considers the application…

Optimization and Control · Mathematics 2026-01-23 Jiachen Jin , Hongxia Wang , Kangkang Deng

We present a novel approach for accelerating AI performance by leveraging Anderson extrapolation, a vector-to-vector mapping technique based on a window of historical iterations. By identifying the crossover point (Fig. 1) where a mixing…

Machine Learning · Computer Science 2024-12-20 Saleem Abdul Fattah Ahmed Al Dajani , David E. Keyes

In this paper we consider the neural network optimization. We develop Anderson-type acceleration method for the stochastic gradient decent method and it improves the network permanence very much. We demonstrate the applicability of the…

Numerical Analysis · Mathematics 2025-12-11 Kazufumi Ito , Tiancheng Xue

We provide rigorous theoretical bounds for Anderson acceleration (AA) that allow for approximate calculations when applied to solve linear problems. We show that, when the approximate calculations satisfy the provided error bounds, the…

Numerical Analysis · Mathematics 2024-04-30 Massimiliano Lupo Pasini , M. Paul Laiu

This work proposes a general strategy for solving possibly nonlinear problems arising from implicit time discretizations as a sequence of explicit solutions. The resulting sequence may exhibit instabilities similar to those of the base…

Numerical Analysis · Mathematics 2025-10-21 Nicolas A. Barnafi , Felipe Galarce , Pablo Brubeck

This paper investigates the use of fixed-point Anderson acceleration method (AA) to a recently proposed hierarchical control framework. Due to its model-free property, the AA-based resulting hierarchical framework becomes more generic since…

Systems and Control · Electrical Eng. & Systems 2021-12-09 Xuan-Huy Pham , Mazen Alamir , François Bonne , Patrick Bonnay

Anderson mixing has been heuristically applied to reinforcement learning (RL) algorithms for accelerating convergence and improving the sampling efficiency of deep RL. Despite its heuristic improvement of convergence, a rigorous…

Machine Learning · Computer Science 2021-10-22 Ke Sun , Yafei Wang , Yi Liu , Yingnan Zhao , Bo Pan , Shangling Jui , Bei Jiang , Linglong Kong

Many computer graphics problems require computing geometric shapes subject to certain constraints. This often results in non-linear and non-convex optimization problems with globally coupled variables, which pose great challenge for…

Graphics · Computer Science 2018-05-16 Yue Peng , Bailin Deng , Juyong Zhang , Fanyu Geng , Wenjie Qin , Ligang Liu

This paper studies the commonly utilized windowed Anderson acceleration (AA) algorithm for fixed-point methods, $x^{(k+1)}=q(x^{(k)})$. It provides the first proof that when the operator $q$ is linear and symmetric the windowed AA, which…

Numerical Analysis · Mathematics 2025-08-01 Casey Garner , Gilad Lerman , Teng Zhang

We propose an Anderson Acceleration (AA) scheme for the adaptive Expectation-Maximization (EM) algorithm for unsupervised learning a finite mixture model from multivariate data (Figueiredo and Jain 2002). The proposed algorithm is able to…

Machine Learning · Computer Science 2020-09-29 Truong Nguyen , Guangye Chen , Luis Chacon

We present a convergence theory for Anderson acceleration (AA) applied to perturbed Newton methods (pNMs) for computing roots of nonlinear problems. Two important special cases are the classical Newton method and the Levenberg-Marquardt…

Numerical Analysis · Mathematics 2025-12-03 Matt Dallas

In this paper we explore acceleration techniques for large scale nonconvex optimization problems with special focuses on deep neural networks. The extrapolation scheme is a classical approach for accelerating stochastic gradient descent for…

Machine Learning · Statistics 2018-05-18 Guangzeng Xie , Yitan Wang , Shuchang Zhou , Zhihua Zhang

In this work, we extend a modified Anderson acceleration proposed in [Y. He, arXiv:2603.25983, 2026] to accelerate the Picard iteration for the Navier-Stokes equations. In this variant of Anderson acceleration, named AAg, the nonlinear…

Numerical Analysis · Mathematics 2026-05-19 Yunhui He , Leo Rebholz