English
Related papers

Related papers: Global Convergence and Error Propagation in Neural…

200 papers

The theory of training deep networks has become a central question of modern machine learning and has inspired many practical advancements. In particular, the gradient descent (GD) optimization algorithm has been extensively studied in…

Optimization and Control · Mathematics 2025-10-29 Alexandru Crăciun , Debarghya Ghoshdastidar

Computational optimal transport (OT) offers a principled framework for generative modeling. Neural OT methods, which use neural networks to learn an OT map (or potential) from data in an amortized way, can be evaluated out of sample after…

Machine Learning · Computer Science 2026-02-04 Alessandro Micheli , Yueqi Cao , Anthea Monod , Samir Bhatt

Gradient Boosting Decision Trees (GBDTs) dominate tabular machine learning, with modern implementations like XGBoost, LightGBM, and CatBoost being based on Newton boosting: a second-order descent step in the space of decision trees. Despite…

Machine Learning · Statistics 2026-05-04 Nikita Zozoulenko , Daniel Falkowski , Thomas Cass , Lukas Gonon

The decentralized gradient descent (DGD) algorithm, and its sibling, diffusion, are workhorses in decentralized machine learning, distributed inference and estimation, and multi-agent coordination. We propose a novel, principled framework…

Signal Processing · Electrical Eng. & Systems 2025-06-04 Erik G. Larsson , Nicolo Michelusi

This work studies the global convergence and implicit bias of Gauss Newton's (GN) when optimizing over-parameterized one-hidden layer networks in the mean-field regime. We first establish a global convergence result for GN in the…

Machine Learning · Computer Science 2023-12-13 Michael Arbel , Romain Menegaux , Pierre Wolinski

In this paper, we study Riemannian zeroth-order optimization in settings where the underlying Riemannian metric $g$ is geodesically incomplete, and the goal is to approximate stationary points with respect to this incomplete metric. To…

Machine Learning · Computer Science 2026-04-14 Shaocong Ma , Heng Huang

This paper considers the problem of decentralized optimization on compact submanifolds, where a finite sum of smooth (possibly non-convex) local functions is minimized by $n$ agents forming an undirected and connected graph. However, the…

Optimization and Control · Mathematics 2025-06-10 Jun Chen , Lina Liu , Tianyi Zhu , Yong Liu , Guang Dai , Yunliang Jiang , Ivor W. Tsang

In recent years, stochastic gradient descent (SGD) based techniques has become the standard tools for training neural networks. However, formal theoretical understanding of why SGD can train neural networks in practice is largely missing.…

Machine Learning · Computer Science 2017-11-03 Yuanzhi Li , Yang Yuan

We consider a class of (possibly strongly) geodesically convex optimization problems on Hadamard manifolds, where the objective function splits into the sum of a smooth and a possibly nonsmooth function. We introduce an intrinsic convex…

Optimization and Control · Mathematics 2025-07-23 Ronny Bergmann , Hajg Jasa , Paula John , Max Pfeffer

We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or…

Machine Learning · Statistics 2020-06-24 Suriya Gunasekar , Jason Lee , Daniel Soudry , Nathan Srebro

We propose a novel evolutionary algorithm for optimizing real-valued objective functions defined on the Grassmann manifold Gr}(k,n), the space of all k-dimensional linear subspaces of R^n. While existing optimization techniques on Gr}(k,n)…

Optimization and Control · Mathematics 2025-03-31 Andrew Lesniewski

When equipped with efficient optimization algorithms, the over-parameterized neural networks have demonstrated high level of performance even though the loss function is non-convex and non-smooth. While many works have been focusing on…

Machine Learning · Computer Science 2021-03-11 Zhiqi Bu , Shiyun Xu , Kan Chen

Graph Neural Networks (GNNs) have demonstrated impressive capabilities in modeling graph-structured data, while Spiking Neural Networks (SNNs) offer high energy efficiency through sparse, event-driven computation. However, existing spiking…

Neural and Evolutionary Computing · Computer Science 2025-08-26 Bowen Zhang , Genan Dai , Hu Huang , Long Lan

We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of shallow neural…

Analysis of PDEs · Mathematics 2026-03-03 Lénaïc Chizat , Maria Colombo , Roberto Colombo , Xavier Fernández-Real

The Trust Region Subproblem is a fundamental optimization problem that takes a pivotal role in Trust Region Methods. However, the problem, and variants of it, also arise in quite a few other applications. In this article, we present a…

Optimization and Control · Mathematics 2022-08-19 Uria Mor , Boris Shustin , Haim Avron

Under the data manifold hypothesis, high-dimensional data are concentrated near a low-dimensional manifold. We study the problem of Riemannian optimization over such manifolds when they are given only implicitly through the data…

Machine Learning · Computer Science 2026-03-03 Andrey Kharitenko , Zebang Shen , Riccardo de Santi , Niao He , Florian Doerfler

Novel convergence analyses are presented of Riemannian stochastic gradient descent (RSGD) on a Hadamard manifold. RSGD is the most basic Riemannian stochastic optimization algorithm and is used in many applications in the field of machine…

Optimization and Control · Mathematics 2023-12-14 Hiroyuki Sakai , Hideaki Iiduka

This paper considers the optimization problem in the form of $\min_{X \in \mathcal{F}_v} f(x) + \lambda \|X\|_1,$ where $f$ is smooth, $\mathcal{F}_v = \{X \in \mathbb{R}^{n \times q} : X^T X = I_q, v \in \mathrm{span}(X)\}$, and $v$ is a…

Optimization and Control · Mathematics 2023-07-21 Wen Huang , Meng Wei , Kyle A. Gallivan , Paul Van Dooren

We consider a distributed non-convex optimization where a network of agents aims at minimizing a global function over the Stiefel manifold. The global function is represented as a finite sum of smooth local functions, where each local…

Optimization and Control · Mathematics 2021-02-16 Shixiang Chen , Alfredo Garcia , Mingyi Hong , Shahin Shahrampour

This paper focuses on minimizing a smooth function combined with a nonsmooth regularization term on a compact Riemannian submanifold embedded in the Euclidean space under a decentralized setting. Typically, there are two types of approaches…

Optimization and Control · Mathematics 2025-07-16 Lei Wang , Le Bao , Xin Liu