English
Related papers

Related papers: A gradient flow on control space with rough initia…

200 papers

Motivated by the occurrence in rate functions of time-dependent large-deviation principles, we study a class of non-negative functions $\mathscr L$ that induce a flow, given by $\mathscr L(\rho_t,\dot\rho_t)=0$. We derive necessary and…

Functional Analysis · Mathematics 2018-01-17 Alexander Mielke , D. R. Michiel Renger , Mark A. Peletier

In deep learning, it is common to use more network parameters than training points. In such scenarioof over-parameterization, there are usually multiple networks that achieve zero training error so that thetraining algorithm induces an…

Machine Learning · Computer Science 2023-08-22 Hung-Hsu Chou , Carsten Gieshoff , Johannes Maly , Holger Rauhut

We study stochastic gradient descent (SGD) for composite optimization problems with $N$ sequential operators subject to perturbations in both the forward and backward passes. Unlike classical analyses that treat gradient noise as additive…

Optimization and Control · Mathematics 2026-02-25 Boao Kong , Hengrui Zhang , Kun Yuan

Gradient-flow analyses show that simplified linear transformers can learn the in-context linear-regression algorithm, but they do not explain the finite-step behavior of gradient descent at large learning rates. Motivated by empirical work…

Machine Learning · Statistics 2026-05-21 Krishnakumar Balasubramanian

Since the early nineties, it has been observed that the Schroedinger bridge problem can be formulated as a stochastic control problem with atypical boundary constraints. This in turn has a fluid dynamic counterpart where the flow of…

Probability · Mathematics 2016-01-20 Yongxin Chen , Tryphon Georgiou , Michele Pavon

We are interested in the gradient flow of a general first order convex functional with respect to the $L^1$-topology. By means of an implicit minimization scheme, we show existence of a global limit solution, which satisfies an…

Analysis of PDEs · Mathematics 2023-10-13 Antonin Chambolle , Matteo Novaga

We consider the problem of finding a saddle point for the convex-concave objective $\min_x \max_y f(x) + \langle Ax, y\rangle - g^*(y)$, where $f$ is a convex function with locally Lipschitz gradient and $g$ is convex and possibly…

Optimization and Control · Mathematics 2021-10-29 Maria-Luiza Vladarean , Yura Malitsky , Volkan Cevher

We propose a variational form of the BDF2 method as an alternative to the commonly used minimizing movement scheme for the time-discrete approximation of gradient flows in abstract metric spaces. Assuming uniform semi-convexity --- but no…

Analysis of PDEs · Mathematics 2017-12-25 Daniel Matthes , Simon Plazotta

We give the first polynomial time algorithms for escaping from high-dimensional saddle points under a moderate number of constraints. Given gradient access to a smooth function $f \colon \mathbb R^d \to \mathbb R$ we show that (noisy)…

Machine Learning · Computer Science 2023-04-21 Dmitrii Avdiukhin , Grigory Yaroslavtsev

We study the convergence of gradient flows related to learning deep linear neural networks (where the activation function is the identity map) from data. In this case, the composition of the network layers amounts to simply multiplying the…

Optimization and Control · Mathematics 2020-10-16 Bubacarr Bah , Holger Rauhut , Ulrich Terstiege , Michael Westdickenberg

Inverse problems in physical or biological sciences often involve recovering an unknown parameter that is random. The sought-after quantity is a probability distribution of the unknown parameter, that produces data that aligns with…

Machine Learning · Statistics 2024-10-02 Qin Li , Maria Oprea , Li Wang , Yunan Yang

Cohen et al. (arXiv:2207.14484) observed that adaptive gradient methods such as Adam operate at the edge of stability. While there has been significant work on continuous-time modeling of gradient descent at the edge of stability, extending…

Machine Learning · Computer Science 2026-05-11 Eric Regis , Sinho Chewi

We study an evolution problem in the space of continuous loops in three-dimensional Euclidean space modelled upon the dynamics of vortex lines in 3d incompressible and inviscid fluids. We establish existence of a local solution starting…

Probability · Mathematics 2007-05-23 Hakima Bessaih , Massimiliano Gubinelli , Francesco Russo

We consider minimizing a nonconvex, smooth function $f$ on a Riemannian manifold $\mathcal{M}$. We show that a perturbed version of Riemannian gradient descent algorithm converges to a second-order stationary point (and hence is able to…

Optimization and Control · Mathematics 2019-06-19 Yue Sun , Nicolas Flammarion , Maryam Fazel

We consider the convex-concave saddle point problem $\min_{x}\max_{y} f(x)+y^\top A x-g(y)$ where $f$ is smooth and convex and $g$ is smooth and strongly convex. We prove that if the coupling matrix $A$ has full column rank, the vanilla…

Optimization and Control · Mathematics 2019-02-05 Simon S. Du , Wei Hu

We analyze the gradient flow of a potential energy in the space of probability measures when we substitute the optimal transport geometry with a geometry based on Sinkhorn divergences, a debiased version of entropic optimal transport. This…

Analysis of PDEs · Mathematics 2025-11-19 Mathis Hardion , Hugo Lavenant

We consider transport over a strongly connected, directed graph. The scheduling amounts to selecting transition probabilities for a discrete-time Markov evolution which is designed to be consistent with certain initial and final marginals.…

Systems and Control · Computer Science 2016-03-29 Yongxin Chen , Tryphon T. Georgiou , Michele Pavon , Allen Tannenbaum

This paper presents a methodology and numerical algorithms for constructing accelerated gradient flows on the space of probability distributions. In particular, we extend the recent variational formulation of accelerated gradient methods in…

Machine Learning · Computer Science 2019-01-14 Amirhossein Taghvaei , Prashant G. Mehta

We analyze the long-time behavior of solutions to semilinear parabolic equations in Euclidean space that arise as gradient flows of an energy functional. We prove that, for general initial data (including data without compact support) the…

Analysis of PDEs · Mathematics 2026-03-03 Daniel Restrepo

Dynamical systems theory has recently been applied in optimization to prove that gradient descent algorithms bypass so-called strict saddle points of the loss function. However, in many modern machine learning applications, the required…

Machine Learning · Computer Science 2024-09-12 Patrick Cheridito , Arnulf Jentzen , Florian Rossmannek