中文
相关论文

相关论文: Global Convergence of Second-order Dynamics in Two…

200 篇论文

We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data. In this case, the composition of the network layers amounts to simply multiplying the…

最优化与控制 · 数学 2020-10-16 Bubacarr Bah , Holger Rauhut , Ulrich Terstiege , Michael Westdickenberg

Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions. For the entropy-regularized RL objective, WPG evolves each…

机器学习 · 计算机科学 2026-05-27 Zhaoyu Zhu , Rui Gao , Shuang Li

The goal of this paper is to discuss some of the results in [31] and [32] and expand upon the work there by proving a global weak existence result as well as a first bubbling analysis in finite time. In addition, an alternative local…

偏微分方程分析 · 数学 2021-12-17 Jerome Wettstein

We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher-student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with…

最优化与控制 · 数学 2026-01-16 Simon Martin , Giulio Biroli , Francis Bach

We study the Wasserstein gradient flow of semi-discrete energies in the space of probability measures, that is functionals depending on two measures-one being an absolutely continuous density and the other an atomic measure. These energies…

偏微分方程分析 · 数学 2026-03-05 Joao Miguel Machado

We continue a long line of research aimed at proving convergence of depth 2 neural networks, trained via gradient descent, to a global minimum. Like in many previous works, our model has the following features: regression with quadratic…

机器学习 · 计算机科学 2022-12-06 Alexander Razborov

In this paper, we study the problem of optimizing a two-layer artificial neural network that best fits a training dataset. We look at this problem in the setting where the number of parameters is greater than the number of sampled points.…

机器学习 · 计算机科学 2017-11-01 Digvijay Boob , Guanghui Lan

The biological plausibility of the backpropagation algorithm has long been doubted by neuroscientists. Two major reasons are that neurons would need to send two different types of signal in the forward and backward phases, and that pairs of…

机器学习 · 计算机科学 2018-08-16 Benjamin Scellier , Anirudh Goyal , Jonathan Binas , Thomas Mesnard , Yoshua Bengio

Natural gradient descent has proven effective at mitigating the effects of pathological curvature in neural network optimization, but little is known theoretically about its convergence properties, especially for \emph{nonlinear} networks.…

机器学习 · 统计学 2019-10-29 Guodong Zhang , James Martens , Roger Grosse

In this article, we introduce a novel concept for second-order information of a nonsmooth function inspired by the Goldstein eps-subdifferential. It comprises the coefficients of all existing second-order Taylor expansions in an eps-ball…

最优化与控制 · 数学 2025-01-06 Bennet Gebken

Gradient-based meta-learning (GBML) with deep neural nets (DNNs) has become a popular approach for few-shot learning. However, due to the non-convexity of DNNs and the bi-level optimization in GBML, the theoretical properties of GBML with…

机器学习 · 计算机科学 2020-11-17 Haoxiang Wang , Ruoyu Sun , Bo Li

Nonsmooth nonconvex optimization problems broadly emerge in machine learning and business decision making, whereas two core challenges impede the development of efficient solution methods with finite-time convergence guarantee: the lack of…

最优化与控制 · 数学 2022-10-18 Tianyi Lin , Zeyu Zheng , Michael I. Jordan

Gradient flows play a substantial role in addressing many machine learning problems. We examine the convergence in continuous-time of a \textit{Fisher-Rao} (Mean-Field Birth-Death) gradient flow in the context of solving convex-concave…

最优化与控制 · 数学 2024-09-19 Razvan-Andrei Lascu , Mateusz B. Majka , Łukasz Szpruch

We extend the standard notion of self-concordance to non-convex optimization and develop a family of second-order algorithms with global convergence guarantees. In particular, two function classes -- \textit{weakly self-concordant}…

最优化与控制 · 数学 2026-04-07 Donald Goldfarb , Lexiao Lai , Tianyi Lin , Jiayu Zhang

The problem of solving partial differential equations (PDEs) can be formulated into a least-squares minimization problem, where neural networks are used to parametrize PDE solutions. A global minimizer corresponds to a neural network that…

数值分析 · 数学 2020-12-14 Tao Luo , Haizhao Yang

We consider a second order gradient flow of the p-elastic energy for a planar theta-network of three curves with fixed lengths. We construct a weak solution of the flow by means of an implicit variational scheme. We show long-time existence…

偏微分方程分析 · 数学 2019-05-24 Matteo Novaga , Paola Pozzi

Large-scale nonconvex optimization problems are ubiquitous in modern machine learning, and among practitioners interested in solving them, Stochastic Gradient Descent (SGD) reigns supreme. We revisit the analysis of SGD in the nonconvex…

最优化与控制 · 数学 2020-07-27 Ahmed Khaled , Peter Richtárik

We analyze Elman-type Recurrent Reural Networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large…

机器学习 · 统计学 2023-03-14 Andrea Agazzi , Jianfeng Lu , Sayan Mukherjee

Flow-based generative models enjoy certain advantages in computing the data generation and the likelihood, and have recently shown competitive empirical performance. Compared to the accumulating theoretical studies on related score-based…

机器学习 · 统计学 2025-06-30 Xiuyuan Cheng , Jianfeng Lu , Yixin Tan , Yao Xie

We study the convergence to equilibrium of the mean field PDE associated with the derivative-free methodologies for solving inverse problems. We show stability estimates in the euclidean Wasserstein distance for the mean field PDE by using…

偏微分方程分析 · 数学 2019-10-24 J. A. Carrillo , U. Vaes