English
Related papers

Related papers: Global Convergence in Training Large-Scale Transfo…

200 papers

We study a system of drift-diffusion PDEs for a potentially infinite number of incompressible phases, subject to a joint pointwise volume constraint. Our analysis is based on the interpretation as a collection of coupled Wasserstein…

Analysis of PDEs · Mathematics 2024-11-22 Clément Cancès , Daniel Matthes , Ismael Medina , Bernhard Schmitzer

Transformer-based models have recently become wildly successful across a diverse set of domains. At the same time, recent work has shown empirically and theoretically that Transformers are inherently limited. Specifically, they argue that…

Machine Learning · Computer Science 2024-07-30 Gbètondji J-S Dovonon , Michael M. Bronstein , Matt J. Kusner

We give a simple local Polyak-Lojasiewicz (PL) criterion that guarantees linear (exponential) convergence of gradient flow and gradient descent to a zero-loss solution of a nonnegative objective. We then verify this criterion for the…

Machine Learning · Computer Science 2026-02-23 Sourav Chatterjee

Many statistical $M$-estimators are based on convex optimization problems formed by the combination of a data-dependent loss function with a norm-based regularizer. We analyze the convergence rates of projected gradient and composite…

Machine Learning · Statistics 2012-07-26 Alekh Agarwal , Sahand N. Negahban , Martin J. Wainwright

Transformers robustly exhibit the ability to perform in-context learning, whereby their predictive accuracy on a task can increase not by parameter updates but merely with the placement of training samples in their context windows. Recent…

Machine Learning · Statistics 2025-10-10 Abhiti Mishra , Yash Patel , Ambuj Tewari

Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds…

Machine Learning · Statistics 2026-05-11 Mana Sakai , Masaaki Imaizumi

We reveal a precise mathematical framework about a new family of generative models which we call Gradient Flow Drifting. With this framework, we prove an equivalence between the recently proposed Drifting Model and the Wasserstein gradient…

Machine Learning · Computer Science 2026-03-12 Jiarui Cao , Zixuan Wei , Yuxin Liu

Wasserstein Gradient Flows (WGF) with respect to specific functionals have been widely used in the machine learning literature. Recently, neural networks have been adopted to approximate certain intractable parts of the underlying…

Machine Learning · Computer Science 2024-01-26 Huminhao Zhu , Fangyikang Wang , Chao Zhang , Hanbin Zhao , Hui Qian

We investigate 1) the rate at which refined properties of the empirical risk---in particular, gradients---converge to their population counterparts in standard non-convex learning tasks, and 2) the consequences of this convergence for…

Machine Learning · Computer Science 2018-11-13 Dylan J. Foster , Ayush Sekhari , Karthik Sridharan

We study random surfaces with a uniformly convex gradient interaction in the presence of quenched disorder taking the form of a random independent external field. Previous work on the model has focused on proving existence and uniqueness of…

Probability · Mathematics 2022-05-09 Paul Dario

We study the JKO scheme for the total variation, characterize the optimizers, prove some of their qualitative properties (in particular a form of maximum principle and in some cases, a minimum principle as well). Finally, we establish a…

Analysis of PDEs · Mathematics 2018-07-09 Guillaume Carlier , Clarice Poon

For free energies of the form \[ F(\mu) = E(\mu) + \sigma\int_\Omega \mu\log\mu\,dx, \quad \sigma > 0, \] we study the Wasserstein gradient flow, a continuity equation also known as mean-field Langevin dynamics, around a stationary state…

Optimization and Control · Mathematics 2026-03-17 Dante Kalise , Lucas M. Moschen , Grigorios A. Pavliotis

We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a…

Machine Learning · Computer Science 2026-04-07 El Mehdi Achour , Kathlén Kohn , Holger Rauhut

Wasserstein distributionally robust optimization (DRO) aims to find robust and generalizable solutions by hedging against data perturbations in Wasserstein distance. Despite its recent empirical success in operations research and machine…

Machine Learning · Computer Science 2022-05-03 Rui Gao

Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning. Here, we study the…

Machine Learning · Computer Science 2022-08-31 Xiaoxia Wu , Edgar Dobriban , Tongzheng Ren , Shanshan Wu , Zhiyuan Li , Suriya Gunasekar , Rachel Ward , Qiang Liu

This paper deals with local criteria for the convergence to a global minimiser for gradient flow trajectories and their discretisations. To obtain quantitative estimates on the speed of convergence, we consider variations on the classical…

Optimization and Control · Mathematics 2024-05-01 Lorenzo Dello Schiavo , Jan Maas , Francesco Pedrotti

This paper considers the problem of solving systems of quadratic equations, namely, recovering an object of interest $\mathbf{x}^{\natural}\in\mathbb{R}^{n}$ from $m$ quadratic equations/samples…

Machine Learning · Statistics 2019-06-13 Yuxin Chen , Yuejie Chi , Jianqing Fan , Cong Ma

This paper generalizes the results obtained by the authors in \cite{dangHomogenizationNondiluteSuspension2021} concerning the homogenization of a non-dilute suspension of magnetic particles in a viscous flow. More specifically, in this…

Analysis of PDEs · Mathematics 2022-02-15 Thuyen Dang , Yuliya Gorb , Silvia Jimenez Bolanos

It has been shown that gradient descent can yield the zero training loss in the over-parametrized regime (the width of the neural networks is much larger than the number of data points). In this work, combining the ideas of some existing…

Optimization and Control · Mathematics 2019-11-05 Lei Li

Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not scalable as they assume access to the complete set of…

Machine Learning · Statistics 2026-03-10 Eduardo Fernandes Montesuma , Yassir Bendou , Mike Gartrell