Related papers: Gradient extremals, talwegs, valleys, and directio…
We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spanned by a few top eigenvectors of the Hessian (equal to the…
Wasserstein gradient flows on probability measures have found a host of applications in various optimization problems. They typically arise as the continuum limit of exchangeable particle systems evolving by some mean-field interaction…
Accelerated gradient descent iterations are widely used in optimization. It is known that, in the continuous-time limit, these iterations converge to a second-order differential equation which we refer to as the accelerated gradient flow.…
Given a non-oscillating gradient trajectory G of a real analytic function f, we show that the limit v of the secants at the limit point O of G along the trajectory G is an eigen-vector of the limit of the direction of the Hessian matrix…
Fractional derivatives are a well-studied generalization of integer order derivatives. Naturally, for optimization, it is of interest to understand the convergence properties of gradient descent using fractional derivatives. Convergence…
Motivated by a constrained minimization problem, it is studied the gradient flows with respect to Hessian Riemannian metrics induced by convex functions of Legendre type. The first result characterizes Hessian Riemannian structures on…
We consider generalized gradients in the general context of $G$-structures. They are natural first order differential operators acting on sections of vector bundles associated to irreducible $G$-representations. We study their geometric…
Gradient descent is the primary workhorse for optimizing large-scale problems in machine learning. However, its performance is highly sensitive to the choice of the learning rate. A key limitation of gradient descent is its lack of natural…
This study leverages the basic insight that the gradient-flow equation associated with the relative Boltzmann entropy, in relation to a Gaussian reference measure within the Hellinger-Kantorovich (HK) geometry, preserves the class of…
Many classical results in algebraic geometry arise from investigating some extremal behaviors that appear among projective varieties not lying on any hypersurface of fixed degree. We study two numerical invariants attached to such…
R. Thom's gradient conjecture states that if a gradient flow of an analytic function converges to a limit, it does so along a unique limiting direction. In this paper, we extend and settle this conjecture in the context of infinite…
This is the first of a series of papers devoted to a thorough analysis of the class of gradient flows in a metric space $(X,\mathsf{d})$ that can be characterized by Evolution Variational Inequalities. We present new results concerning the…
The logarithmic divergence is an extension of the Bregman divergence motivated by optimal transport and a generalized convex duality, and satisfies many remarkable properties. Using the geometry induced by the logarithmic divergence, we…
Propagation and tunneling of light through subwavelength photonic barriers, formed by dielectric layers with continuous spatial variations of dielectric susceptibility across the film are considered. Effects of giant heterogeneity-induced…
We give inequalities relating the eigenvalues of the adjacency matrix and the Laplacian of a graph, and its minimum and maximum degrees. The results are applied to derive new conditions for quasi-randomness of graphs.
In [8], the gradient conjecture of R. Thom was proven for gradient flows of analytic functions on Rn. This result means that the secant at a limit point converges, so that the flow cannot spiral forever. Once the trajectory becomes…
Stochastic gradients for deep neural networks exhibit strong correlations along the optimization trajectory, and are often aligned with a small set of Hessian eigenvectors associated with outlier eigenvalues. Recent work shows that…
Finding latent structures in data is drawing increasing attention in diverse fields such as image and signal processing, fluid dynamics, and machine learning. In this work we examine the problem of finding the main modes of gradient flows.…
The gradient-flow dynamics of an arbitrary geometric quantity is derived using a generalization of Darcy's Law. We consider flows in both Lagrangian and Eulerian formulations. The Lagrangian formulation includes a dissipative modification…
We consider the set of extremal points of the generalized unit ball induced by gradient total variation seminorms for vector-valued functions on bounded Euclidean domains. These are central to the understanding of sparse solutions and…