Related papers: Global Convergence of Second-order Dynamics in Two…
We study a continuous-time dynamical system which arises as the limit of a broad class of nonlinearly preconditioned gradient methods. Under mild assumptions, we establish existence of global solutions and derive Lyapunov-based convergence…
We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of shallow neural…
We develop novel neural network-based implicit particle methods to compute high-dimensional Wasserstein-type gradient flows with linear and nonlinear mobility functions. The main idea is to use the Lagrangian formulation in the…
Motivated by a constrained minimization problem, it is studied the gradient flows with respect to Hessian Riemannian metrics induced by convex functions of Legendre type. The first result characterizes Hessian Riemannian structures on…
Deep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in deep neural network (DNN) training. However, theoretical…
We prove an existence result for a large class of PDEs with a nonlinear Wasserstein gradient flow structure. We use the classical theory of Wasserstein gradient flow to derive an EDI formulation of our PDE and prove that under some…
This paper contributes to the exploration of a recently introduced computational paradigm known as second-order flows, which are characterized by novel dissipative hyperbolic partial differential equations extending accelerated gradient…
We propose a variational finite volume scheme to approximate the solutions to Wasserstein gradient flows. The time discretization is based on an implicit linearization of the Wasserstein distance expressed thanks to Benamou-Brenier formula,…
Despite recent theoretical progress on the non-convex optimization of two-layer neural networks, it is still an open question whether gradient descent on neural networks without unnatural modifications can achieve better sample complexity…
A key challenge in modern deep learning theory is to explain the remarkable success of gradient-based optimization methods when training large-scale, complex deep neural networks. Though linear convergence of such methods has been proved…
A fully coupled system of two second-order parabolic degenerate equations arising as a thin film approximation to the Muskat problem is interpreted as a gradient flow for the 2-Wasserstein distance in the space of probability measures with…
We prove the equivalence between the notion of Wasserstein gradient flow for a one-dimensional nonlocal transport PDE with attractive/repulsive Newtonian potential on one side, and the notion of entropy solution of a Burgers-type scalar…
Recent studies of gradient descent with large step sizes have shown that there is often a regime with an initial increase in the largest eigenvalue of the loss Hessian (progressive sharpening), followed by a stabilization of the eigenvalue…
We provide a numerical analysis and computation of neural network projected schemes for approximating one dimensional Wasserstein gradient flows. We approximate the Lagrangian mapping functions of gradient flows by the class of two-layer…
We consider the scenario of supervised learning in Deep Learning (DL) networks, and exploit the arbitrariness of choice in the Riemannian metric relative to which the gradient descent flow can be defined (a general fact of differential…
We provide an estimation of the dissipation of the Wasserstein 2 distance between the law of some interacting $N$-particle system, and the $N$ times tensorized product of solution to the corresponding limit nonlinear conservation law. It…
We study the 2D Ginzburg-Landau theory for a type-II superconductor in an applied magnetic field varying between the second and third critical value. In this regime the order parameter minimizing the GL energy is concentrated along the…
The mean field (MF) theory of multilayer neural networks centers around a particular infinite-width scaling, where the learning dynamics is closely tracked by the MF limit. A random fluctuation around this infinite-width limit is expected…
We propose a generalization of the gradient flow equation for quantum field theories with nonlinearly realized symmetry. Applying the equation to $\mathcal{N}=1$ $SU(N)$ super Yang-Mills theory in four dimensions, we construct a…
Current state-of-the-art analyses on the convergence of gradient descent for training neural networks focus on characterizing properties of the loss landscape, such as the Polyak-Lojaciewicz (PL) condition and the restricted strong…