相关论文: On Hessian limit directions along non-oscillating …
Let (X,O) be a real analytic isolated surface singularity at the origin o of a real analytic manifold M equipped with a real analytic metric g. Given a real analytic function f:(M,O) --> (R,0) singular at O, we prove that the gradient…
Gradient extremals are loci along which the gradient is an eigenvector of the Hessian. These objects provide a natural geometric framework connecting several notions, notably valleys and talwegs, which we analyze from a variational…
It is well known that for analytic cost functions, gradient flow trajectories have finite length and converge to a single critical point. The gradient conjecture of R. Thom states that, again for analytic cost functions, whenever the…
We study stochastic zeroth order gradient and Hessian estimators for real-valued functions in $\mathbb{R}^n$. We show that, via taking finite difference along random orthogonal directions, the variance of the stochastic finite difference…
Let x(t) be a trajectory of the gradient of a real analytic function and suppose that x_0 is a limit point of x(t). We prove the gradient conjecture of R. Thom which states that the secants of x(t) at x_0 have a limit. Actually we show a…
Let G be a connected real reductive Lie group acting linearly on a finite dimensional vector space V over R. This action admits a Kempf-Ness function and so we have an associated gradient map. If G is Abelian we explicitly compute the image…
We construct an example of a real plane analytic singular metric, degenerating only at the origin, such that any gradient trajectory (respectively to this singular metric) of some well chosen function spirals around the origin. The…
Fractional derivatives are a well-studied generalization of integer order derivatives. Naturally, for optimization, it is of interest to understand the convergence properties of gradient descent using fractional derivatives. Convergence…
Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the Hessian…
We consider large non-Hermitian random matrices $X$ with complex, independent, identically distributed centred entries and show that the linear statistics of their eigenvalues are asymptotically Gaussian for test functions having…
An unbiased low-variance gradient estimator, termed GO gradient, was proposed recently for expectation-based objectives $\mathbb{E}_{q_{\boldsymbol{\gamma}}(\boldsymbol{y})} [f(\boldsymbol{y})]$, where the random variable (RV)…
Consider sample covariance matrices of the form $Q:=\Sigma^{1/2} X X^\top \Sigma^{1/2}$, where $X=(x_{ij})$ is an $n\times N$ random matrix whose entries are independent random variables with mean zero and variance $N^{-1}$, and $\Sigma$ is…
The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold $2/\eta$ during gradient descent (GD) with step size $\eta$. Despite (apparently)…
Operator self-similar processes, as an extension of self-similar processes, have been studied extensively. In this work, we study limit theorems for functionals of Gaussian vectors. Under some conditions, we determine that the limit of…
Recent research has observed that in machine learning optimization, gradient descent (GD) often operates at the edge of stability (EoS) [Cohen, et al., 2021], where the stepsizes are set to be large, resulting in non-monotonic losses…
Consider a dominant rational self-map $f$ on a smooth projective variety $X$ defined over $\overline{\mathbb{Q}}$. We prove that \begin{align} \lim_{n \to \infty} \frac{h_{Y}(f^{n}(x))}{h_{H}(f^{n}(x)) } = 0, \end{align} where $h_{Y}$ is a…
Let $\Delta\subsetneq \mathbb C$ be a simply connected domain, let $f:\mathbb D \to \Delta$ be a Riemann map and let $\{z_k\}\subset \Delta$ be a compactly divergent sequence. Using Gromov's hyperbolicity theory, we show that…
We study the scaling limit of statistical mechanics models with non-convex Hamiltonians that are gradient perturbations of Gaussian measures. Characterising features of our gradient models are the imposed boundary tilt and the surface…
We develop optimization methods which offer new trade-offs between the number of gradient and Hessian computations needed to compute the critical point of a non-convex function. We provide a method that for any twice-differentiable $f\colon…
In a real Hilbert space setting, we study the convergence properties of an inexact gradient algorithm featuring both viscous and Hessian driven damping for convex differentiable optimization. In this algorithm, the gradient evaluation can…