中文
相关论文

相关论文: Global Convergence of Second-order Dynamics in Two…

200 篇论文

We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence…

最优化与控制 · 数学 2026-04-20 Konstantinos Oikonomidis , Alexander Bodard , Jan Quan , Panagiotis Patrinos

We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of shallow neural…

偏微分方程分析 · 数学 2026-03-03 Lénaïc Chizat , Maria Colombo , Roberto Colombo , Xavier Fernández-Real

We develop novel neural network-based implicit particle methods to compute high-dimensional Wasserstein-type gradient flows with linear and nonlinear mobility functions. The main idea is to use the Lagrangian formulation in the…

数值分析 · 数学 2023-11-14 Wonjun Lee , Li Wang , Wuchen Li

Motivated by a constrained minimization problem, it is studied the gradient flows with respect to Hessian Riemannian metrics induced by convex functions of Legendre type. The first result characterizes Hessian Riemannian structures on…

最优化与控制 · 数学 2018-11-27 Felipe Alvarez , Jérôme Bolte , Olivier Brahic

Deep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in deep neural network (DNN) training. However, theoretical…

最优化与控制 · 数学 2019-05-14 Jinshan Zeng , Tim Tsz-Kit Lau , Shaobo Lin , Yuan Yao

We prove an existence result for a large class of PDEs with a nonlinear Wasserstein gradient flow structure. We use the classical theory of Wasserstein gradient flow to derive an EDI formulation of our PDE and prove that under some…

偏微分方程分析 · 数学 2024-07-31 Thibault Caillet , Filippo Santambrogio

This paper contributes to the exploration of a recently introduced computational paradigm known as second-order flows, which are characterized by novel dissipative hyperbolic partial differential equations extending accelerated gradient…

数值分析 · 数学 2025-05-13 Haifan Chen , Guozhi Dong , José A. Iglesias , Wei Liu , Ziqing Xie

We propose a variational finite volume scheme to approximate the solutions to Wasserstein gradient flows. The time discretization is based on an implicit linearization of the Wasserstein distance expressed thanks to Benamou-Brenier formula,…

数值分析 · 数学 2019-07-22 Clément Cancès , Thomas O. Gallouët , Gabriele Todeschi

Despite recent theoretical progress on the non-convex optimization of two-layer neural networks, it is still an open question whether gradient descent on neural networks without unnatural modifications can achieve better sample complexity…

机器学习 · 计算机科学 2023-10-10 Arvind Mahankali , Jeff Z. Haochen , Kefan Dong , Margalit Glasgow , Tengyu Ma

A key challenge in modern deep learning theory is to explain the remarkable success of gradient-based optimization methods when training large-scale, complex deep neural networks. Though linear convergence of such methods has been proved…

机器学习 · 计算机科学 2025-09-30 Yash Jakhmola

A fully coupled system of two second-order parabolic degenerate equations arising as a thin film approximation to the Muskat problem is interpreted as a gradient flow for the 2-Wasserstein distance in the space of probability measures with…

偏微分方程分析 · 数学 2013-08-29 Philippe Laurencot , Bogdan-Vasile Matioc

We prove the equivalence between the notion of Wasserstein gradient flow for a one-dimensional nonlocal transport PDE with attractive/repulsive Newtonian potential on one side, and the notion of entropy solution of a Burgers-type scalar…

偏微分方程分析 · 数学 2013-10-16 Giovanni A. Bonaschi , José A. Carrillo , Marco Di Francesco , Mark A. Peletier

Recent studies of gradient descent with large step sizes have shown that there is often a regime with an initial increase in the largest eigenvalue of the loss Hessian (progressive sharpening), followed by a stabilization of the eigenvalue…

机器学习 · 计算机科学 2022-10-11 Atish Agarwala , Fabian Pedregosa , Jeffrey Pennington

We provide a numerical analysis and computation of neural network projected schemes for approximating one dimensional Wasserstein gradient flows. We approximate the Lagrangian mapping functions of gradient flows by the class of two-layer…

数值分析 · 数学 2024-02-27 Xinzhe Zuo , Jiaxi Zhao , Shu Liu , Stanley Osher , Wuchen Li

We consider the scenario of supervised learning in Deep Learning (DL) networks, and exploit the arbitrariness of choice in the Riemannian metric relative to which the gradient descent flow can be defined (a general fact of differential…

机器学习 · 计算机科学 2026-05-26 Thomas Chen

We provide an estimation of the dissipation of the Wasserstein 2 distance between the law of some interacting $N$-particle system, and the $N$ times tensorized product of solution to the corresponding limit nonlinear conservation law. It…

偏微分方程分析 · 数学 2018-10-23 Samir Salem

We study the 2D Ginzburg-Landau theory for a type-II superconductor in an applied magnetic field varying between the second and third critical value. In this regime the order parameter minimizing the GL energy is concentrated along the…

数学物理 · 物理学 2019-10-01 M. Correggi , N. Rougerie

The mean field (MF) theory of multilayer neural networks centers around a particular infinite-width scaling, where the learning dynamics is closely tracked by the MF limit. A random fluctuation around this infinite-width limit is expected…

机器学习 · 计算机科学 2021-11-01 Huy Tuan Pham , Phan-Minh Nguyen

We propose a generalization of the gradient flow equation for quantum field theories with nonlinearly realized symmetry. Applying the equation to $\mathcal{N}=1$ $SU(N)$ super Yang-Mills theory in four dimensions, we construct a…

高能物理 - 格点 · 物理学 2015-11-23 Sinya Aoki , Kengo Kikuchi , Tetsuya Onogi

Current state-of-the-art analyses on the convergence of gradient descent for training neural networks focus on characterizing properties of the loss landscape, such as the Polyak-Lojaciewicz (PL) condition and the restricted strong…

机器学习 · 计算机科学 2024-01-08 Fangshuo Liao , Anastasios Kyrillidis