Related papers: Gradient extremals, talwegs, valleys, and directio…
This paper considers the problem of solving systems of quadratic equations, namely, recovering an object of interest $\mathbf{x}^{\natural}\in\mathbb{R}^{n}$ from $m$ quadratic equations/samples…
We are interested in existence of gradient flows for shape functionals especially for first Laplacian eigenvalues. We introduce different techniques to prove existence and use different formulations for gradient flows. We apply a…
Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that in the early stages of training, the weights remain small and near the origin, but…
In a Hilbert setting, for convex differentiable optimization, we develop a general framework for adaptive accelerated gradient methods. They are based on damped inertial dynamics where the coefficients are designed in a closed-loop way.…
We rigorously study the relation between the training dynamics via stochastic gradient descent (SGD) and the spectra of empirical Hessian and gradient matrices. We prove that in two canonical classification tasks for multi-class…
We associate curves of isotropic, Lagrangian and coisotropic subspaces to higher order, one parameter variational problems. Minimality and conjugacy properties of extremals are described in terms of these curves.
In the theory of so called "Covariant Quantum Mechanics" a basic role is played by Hermitian vector fields on a complex line bundle in the frameworks of Galilei and Einstein spacetimes. In fact, it has been proved that the Lie algebra of…
Gradient descent-ascent (GDA) flows play a central role in finding saddle points of bivariate functionals, with applications in optimization, game theory, and robust control. While they are well-understood in Hilbert and Banach spaces via…
Various questions related to distances between vertices of simple, finite graphs are of interest to extremal graph theorists. The Steiner distance of a set of $k$ vertices is a natural generalization of the regular distance. We extend…
Full-batch gradient descent on neural networks drives the largest Hessian eigenvalue to the threshold $2/\eta$, where $\eta$ is the learning rate. This phenomenon, the Edge of Stability, has resisted a unified explanation: existing accounts…
The Hermitian eigenvalue problem asks for the possible eigenvalues of a sum of $n\times n$ Hermitian matrices, given the eigenvalues of the summands. The regular faces of the cones $\Gamma_n(s)$ controlling this problem have been…
We study the learning performance of gradient descent when the empirical risk is weakly convex, namely, the smallest negative eigenvalue of the empirical risk's Hessian is bounded in magnitude. By showing that this eigenvalue can control…
We investigate two-point velocity-gradient correlation functions in homogeneous isotropic turbulence using exact relations and direct numerical simulations. The second-order gradient correlation is shown to be exactly related to the…
We empirically demonstrate that full-batch gradient descent on neural network training objectives typically operates in a regime we call the Edge of Stability. In this regime, the maximum eigenvalue of the training loss Hessian hovers just…
This work revolves around the rigorous asymptotic analysis of models in nonlocal hyperelasticity. The corresponding variational problems involve integral functionals depending on nonlocal gradients with a finite interaction range $\delta$,…
This article deals with the conjugate gradient method on a Riemannian manifold with interest in global convergence analysis. The existing conjugate gradient algorithms on a manifold endowed with a vector transport need the assumption that…
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…
We prove the convergence of a Wasserstein gradient flow of a free energy in inhomogeneous media. Both the energy and media can depend on the spatial variable in a fast oscillatory manner. In particular, we show that the gradient-flow…
We prove a central limit theorem for the components of the largest eigenvectors of the adjacency matrix of a finite-dimensional random dot product graph whose true latent positions are unknown. In particular, we follow the methodology…
Wasserstein gradient flows have become a central tool for optimization problems over probability measures. A natural numerical approach is forward-Euler time discretization. We show, however, that even in the simple case where the energy…