中文
相关论文

相关论文: Global Convergence of Second-order Dynamics in Two…

200 篇论文

The paper surveys recent progresses in understanding the dynamics and loss landscape of the gradient flow equations associated to deep linear neural networks, i.e., the gradient descent training dynamics (in the limit when the step size…

机器学习 · 计算机科学 2025-11-14 Joel Wendin , Claudio Altafini

We prove the optimal $W^{2,\infty}$ regularity for variational problems with convex gradient constraints. We do not assume any regularity of the constraints; so the constraints can be nonsmooth, and they need not be strictly convex. When…

偏微分方程分析 · 数学 2021-01-27 Mohammad Safdari

We study whether a depth two neural network can learn another depth two network using gradient descent. Assuming a linear output node, we show that the question of whether gradient descent converges to the target function is equivalent to…

数据结构与算法 · 计算机科学 2018-12-06 Rina Panigrahy , Sushant Sachdeva , Qiuyi Zhang

Wasserstein gradient flows provide a powerful means of understanding and solving many diffusion equations. Specifically, Fokker-Planck equations, which model the diffusion of probability measures, can be understood as gradient descent over…

机器学习 · 计算机科学 2021-10-26 Petr Mokrov , Alexander Korotin , Lingxiao Li , Aude Genevay , Justin Solomon , Evgeny Burnaev

In this paper, we theoretically prove that gradient descent can find a global minimum of non-convex optimization of all layers for nonlinear deep neural networks of sizes commonly encountered in practice. The theory developed in this paper…

机器学习 · 统计学 2020-06-18 Kenji Kawaguchi , Jiaoyang Huang

We develop a geometric convergence theory for neural-network optimization within the minimizing movement scheme (MMS) framework. Reformulating each neural MMS step as a minimization over the set of increments in a Hilbert space, we show…

最优化与控制 · 数学 2026-05-28 Shixin Zheng , Yiwei Wang , Haizhao Yang

Neural networks trained to minimize the logistic (a.k.a. cross-entropy) loss with gradient-based methods are observed to perform well in many supervised classification tasks. Towards understanding this phenomenon, we analyze the training…

最优化与控制 · 数学 2020-06-23 Lenaic Chizat , Francis Bach

Graph Neural Networks (GNNs) are powerful tools for addressing learning problems on graph structures, with a wide range of applications in molecular biology and social networks. However, the theoretical foundations underlying their…

机器学习 · 计算机科学 2025-01-27 Dhiraj Patel , Anton Savostianov , Michael T. Schaub

In this paper, we attempt to compare two distinct branches of research on second-order optimization methods. The first one studies self-concordant functions and barriers, the main assumption being that the third derivative of the objective…

最优化与控制 · 数学 2024-08-21 Pavel Dvurechensky , Yurii Nesterov

We study first order methods to compute the barycenter of a probability distribution $P$ over the space of probability measures with finite second moment. We develop a framework to derive global rates of convergence for both gradient…

统计理论 · 数学 2020-06-16 Sinho Chewi , Tyler Maunu , Philippe Rigollet , Austin J. Stromme

The second-order hydrodynamic equations for evolution of shear and bulk viscous pressure have been derived within the framework of covariant kinetic theory based on the effective fugacity quasiparticle model. The temperature-dependent…

高能物理 - 唯象学 · 物理学 2021-09-22 Samapan Bhadury , Manu Kurian , Vinod Chandra , Amaresh Jaiswal

A deep equilibrium model uses implicit layers, which are implicitly defined through an equilibrium point of an infinite sequence of computation. It avoids any explicit computation of the infinite sequence by finding an equilibrium point…

机器学习 · 计算机科学 2021-02-19 Kenji Kawaguchi

Gradient flow in the 2-Wasserstein space is widely used to optimize functionals over probability distributions and is typically implemented using an interacting particle system with $n$ particles. Analyzing these algorithms requires showing…

机器学习 · 计算机科学 2026-03-27 Chandan Tankala , Dheeraj M. Nagaraj , Anant Raj

We study the discretization of generalized Wasserstein distances with nonlinear mobilities on the real line via suitable discrete metrics on the cone of N ordered particles, a setting which naturally appears in the framework of…

偏微分方程分析 · 数学 2022-09-01 Simone Di Marino , Lorenzo Portinale , Emanuela Radici

We study the stochastic optimization problem from a continuous-time perspective, with a focus on the Stochastic Gradient Descent with Momentum (SGDM) method. We show that the trajectory of SGDM, despite its \emph{stochastic} nature,…

最优化与控制 · 数学 2025-07-17 Yasong Feng , Yifan Jiang , Tianyu Wang , Zhiliang Ying

We study the gradient descent (GD) dynamics of a depth-2 linear neural network with a single input and output. We show that GD converges at an explicit linear rate to a global minimum of the training loss, even with a large stepsize --…

机器学习 · 计算机科学 2025-01-22 Pierfrancesco Beneventano , Blake Woodworth

In this work we study the asymptotic behavior of a class of damped second-order gradient systems $$ \ddot{u}(t) + a\dot{u}(t) + \nabla W(u(t)) = 0, $$ under assumptions ensuring local convexity of the potential near equilibrium and…

经典分析与常微分方程 · 数学 2025-12-25 Renan J. S. Isneri , Eric B. Santiago , Severino H. da Silva

We show that the continuous-time gradient descent in Rn can be viewed as an optimal controlled evolution for a suitable action functional; a similar result holds for stochastic gradient descent. We then provide an analogous characterization…

最优化与控制 · 数学 2025-11-03 Yongxin Chen , Tryphon Georgiou , Michele Pavon

Overparametrization is a key factor in the absence of convexity to explain global convergence of gradient descent (GD) for neural networks. Beside the well studied lazy regime, infinite width (mean field) analysis has been developed for…

神经与进化计算 · 计算机科学 2023-02-07 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

In this paper, we study a consensus-based optimization method for nonconvex bi-level optimization, where the objective is to minimize an upper-level function over the set of global minimizers of a lower-level problem. The proposed approach…

最优化与控制 · 数学 2026-05-20 Yutong Chao , Xudong Sun , Konstantin Riedl , Majid Khadiv , Jalal Etesami