Related papers: Gradient extremals, talwegs, valleys, and directio…
We derive new gradient flows of divergence functions in the probability space embedded with a class of Riemannian metrics. The Riemannian metric tensor is built from the transported Hessian operator of an entropy function. The new gradient…
The convergence theory for the gradient sampling algorithm is extended to directionally Lipschitz functions. Although directionally Lipschitz functions are not necessarily locally Lipschitz, they are almost everywhere differentiable and…
A new formulation of boundary value problems in gradient elasticity is presented in this work. The main outcome is the construction of partial differential systems of second order, which are typically equivalent with the well known fourth…
This paper addresses the question of when projections of a high-dimensional random vector are approximately Gaussian. This problem has been studied previously in the context of high-dimensional data analysis, where the focus is on…
Stochastic coordinate descent algorithms are efficient methods in which each iterate is obtained by fixing most coordinates at their values from the current iteration, and approximately minimizing the objective with respect to the remaining…
The study of first-order optimization is sensitive to the assumptions made on the objective functions. These assumptions induce complexity classes which play a key role in worst-case analysis, including the fundamental concept of algorithm…
Traditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(\theta)$, is bounded by $2/\eta$, training is "stable" and the training loss decreases monotonically. Recent…
We propose a categorical semantics of gradient-based machine learning algorithms in terms of lenses, parametrised maps, and reverse derivative categories. This foundation provides a powerful explanatory and unifying framework: it…
Classical analyses of gradient descent (GD) define a stability threshold based on the largest eigenvalue of the loss Hessian, often termed sharpness. When the learning rate lies below this threshold, training is stable and the loss…
In this paper, we study Hessian equations with prescribed contact angle boundary value or oblique derivative boundary value and finally derive the a priori global gradient estimate for the admissible solutions.
Wave transport in a media with slow spatial gradient of its characteristics is found to exhibit a universal wave pattern ("gradient marker") in a vicinity of the maxima/minima of the gradient. The pattern is common for optics, quantum…
We consider inhomogeneous Erd\H{o}s-R\'enyi graphs. We suppose that the maximal mean degree $d$ satisfies $d \ll \log n$. We characterize the asymptotic behavior of the $n^{1 - o(1)}$ largest eigenvalues of the adjacency matrix and its…
We present a primal only derivation of Mirror Descent as a "partial" discretization of gradient flow on a Riemannian manifold where the metric tensor is the Hessian of the Mirror Descent potential. We contrast this discretization to Natural…
We prove gradient estimates for hypersurfaces in the hyperbolic space $\mathbb{H}^{n+1},$ expanding by negative powers of a certain class of homogeneous curvature functions. We obtain optimal gradient estimates for hypersurfaces evolving by…
We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a strong component along these directions. Furthermore, we show…
We consider an evolution equation similar to that introduced by Vese and whose solution converges in large time to the convex envelope of the initial datum. We give a stochastic control representation for the solution from which we deduce,…
We suggest simple implementable modifications of conditional gradient and gradient projection methods for smooth convex optimization problems in Hilbert spaces. Usually, the custom methods attain only weak convergence. We prove strong…
We consider a class of sparse random matrices, which includes the adjacency matrix of Erd\H{o}s-R\'enyi graph ${\bf G}(N,p)$. For $N^{-1+o(1)}\leq p\leq 1/2$, we show that the non-trivial edge eigenvectors are asymptotically jointly normal.…
Gravity currents are a ubiquitous density driven flow occurring in both the natural environment and in industry. They include: seafloor turbidity currents, primary vectors of sediment, nutrient and pollutant transport; cold fronts; and…
This letter is dedicated to providing proof of two statements concerning the gradient expansion of relativistic hydrodynamics. The first statement is that \textit{the ordering of transverse derivatives is irrelevant in the gradient…