中文
相关论文

相关论文: Global Convergence of Second-order Dynamics in Two…

200 篇论文

We study the convergence of gradient flow for the training of deep neural networks. If Residual Neural Networks are a popular example of very deep architectures, their training constitutes a challenging optimization problem due notably to…

机器学习 · 计算机科学 2025-07-22 Raphaël Barboni , Gabriel Peyré , François-Xavier Vialard

Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in…

机器学习 · 统计学 2024-11-01 Cheng Gao , Yuan Cao , Zihao Li , Yihan He , Mengdi Wang , Han Liu , Jason Matthew Klusowski , Jianqing Fan

Finding the optimal configuration of parameters in ResNet is a nonconvex minimization problem, but first-order methods nevertheless find the global optimum in the overparameterized regime. We study this phenomenon with mean-field analysis,…

机器学习 · 计算机科学 2021-11-30 Zhiyan Ding , Shi Chen , Qin Li , Stephen Wright

In the mean field regime, neural networks are appropriately scaled so that as the width tends to infinity, the learning dynamics tends to a nonlinear and nontrivial dynamical limit, known as the mean field limit. This lends a way to study…

机器学习 · 计算机科学 2021-05-12 Huy Tuan Pham , Phan-Minh Nguyen

We study the training dynamics of shallow neural networks, in a two-timescale regime in which the stepsizes for the inner layer are much smaller than those for the outer layer. In this regime, we prove convergence of the gradient flow to a…

最优化与控制 · 数学 2023-10-26 Pierre Marion , Raphaël Berthier

The generalization mystery of overparametrized deep nets has motivated efforts to understand how gradient descent (GD) converges to low-loss solutions that generalize well. Real-life neural networks are initialized from small random values…

机器学习 · 计算机科学 2021-11-10 Kaifeng Lyu , Zhiyuan Li , Runzhe Wang , Sanjeev Arora

Many tasks in machine learning and signal processing can be solved by minimizing a convex function of a measure. This includes sparse spikes deconvolution or training a neural network with a single hidden layer. For these problems, we study…

最优化与控制 · 数学 2018-10-30 Lenaic Chizat , Francis Bach

This paper investigates the gradient flow structure, well-posedness, and asymptotic behavior of the Fokker-Planck equation defined on locally uniformly finite graphs, which is highly non-trivial compared with the finite case. We first…

概率论 · 数学 2025-11-13 Cong Wang

Our work is motivated by a desire to study the theoretical underpinning for the convergence of stochastic gradient type algorithms widely used for non-convex learning tasks such as training of neural networks. The key insight, already…

概率论 · 数学 2020-12-15 Kaitong Hu , Zhenjie Ren , David Siska , Lukasz Szpruch

In this paper, we study higher-order-accurate-in-time minimizing movements schemes for Wasserstein gradient flows. We introduce a novel accelerated second-order scheme, leveraging the differential structure of the Wasserstein space in both…

偏微分方程分析 · 数学 2025-12-23 Raymond Chu , Matt Jacobs

Fitting a function by using linear combinations of a large number $N$ of `simple' components is one of the most fruitful ideas in statistical learning. This idea lies at the core of a variety of methods, from two-layer neural networks to…

统计理论 · 数学 2019-08-20 Adel Javanmard , Marco Mondelli , Andrea Montanari

In a recent work, we introduced a rigorous framework to describe the mean field limit of the gradient-based learning dynamics of multilayer neural networks, based on the idea of a neuronal embedding. There we also proved a global…

机器学习 · 计算机科学 2020-06-17 Huy Tuan Pham , Phan-Minh Nguyen

This paper studies minimax optimization problems defined over infinite-dimensional function classes of overparameterized two-layer neural networks. In particular, we consider the minimax optimization problem stemming from estimating linear…

机器学习 · 计算机科学 2024-10-25 Yuchen Zhu , Yufeng Zhang , Zhaoran Wang , Zhuoran Yang , Xiaohong Chen

We study the optimization of wide neural networks (NNs) via gradient flow (GF) in setups that allow feature learning while admitting non-asymptotic global convergence guarantees. First, for wide shallow NNs under the mean-field scaling and…

机器学习 · 计算机科学 2022-04-25 Zhengdao Chen , Eric Vanden-Eijnden , Joan Bruna

Physics informed neural networks (PINNs) represent a very popular class of neural solvers for partial differential equations. In practice, one often employs stochastic gradient descent type algorithms to train the neural network. Therefore,…

机器学习 · 计算机科学 2025-09-01 Bangti Jin , Longjun Wu

We study existence and long-time behavior of weak solutions to a thin-film equation with a confinement potential and a second-order degenerate diffusion term. It is known that in absence of second order effects, solutions for general…

偏微分方程分析 · 数学 2025-05-14 Christian Parsch

In this work, we consider smooth unconstrained optimization problems and we deal with the class of gradient methods with momentum, i.e., descent algorithms where the search direction is defined as a linear combination of the current…

最优化与控制 · 数学 2025-12-04 Matteo Lapucci , Giampaolo Liuzzi , Stefano Lucidi , Davide Pucci , Marco Sciandrone

The training dynamics of two-layer neural networks with batch normalization (BN) is studied. It is written as the training dynamics of a neural network without BN on a Riemannian manifold. Therefore, we identify BN's effect of changing the…

机器学习 · 计算机科学 2021-10-19 Chao Ma , Lexing Ying

Neural networks with a large number of parameters admit a mean-field description, which has recently served as a theoretical explanation for the favorable training properties of "overparameterized" models. In this regime, gradient descent…

机器学习 · 统计学 2019-03-28 Grant Rotskoff , Samy Jelassi , Joan Bruna , Eric Vanden-Eijnden

This study focuses on a Wasserstein-type gradient flow, which represents an optimization process of a continuous model of a Deep Neural Network (DNN). First, we establish the existence of a minimizer for an average loss of the model under…

机器学习 · 计算机科学 2024-04-16 Noboru Isobe
‹ 上一页 1 2 3 10 下一页 ›