Related papers: Gradient extremals, talwegs, valleys, and directio…
Extremal graphical models encode the conditional independence structure of multivariate extremes. Key statistics for learning extremal graphical structures are empirical extremal variograms, for which we prove non-asymptotic concentration…
Motivated by a wide variety of applications, ranging from stochastic optimization to dimension reduction through variable selection, the problem of estimating gradients accurately is of crucial importance in statistics and learning theory.…
The analysis of gradient descent-type methods typically relies on the Lipschitz continuity of the objective gradient. This generally requires an expensive hyperparameter tuning process to appropriately calibrate a stepsize for a given…
General equations are derived for slow viscous thin fluid film flows on curved surfaces through an extension of Leal's pedagogical approach, which leaves the characteristic velocity scale unspecified and employs a direct through-thickness…
We consider the dynamics of vector fields on three-manifolds which are constrained to lie within a plane field, such as occurs in nonholonomic dynamics. On compact manifolds, such vector fields force dynamics beyond that of a gradient flow,…
We present a new accelerated stochastic second-order method that is robust to both gradient and Hessian inexactness, which occurs typically in machine learning. We establish theoretical lower bounds and prove that our algorithm achieves…
The paper is devoted to multidimensional $(0,1)$-matrices extremal with respect to containing a polydiagonal (a fractional generalization of a diagonal). Every extremal matrix is a threshold matrix, i.e., an entry belongs to its support…
Curves in Lagrange Grassmannians naturally appear when one studies intrinsically "the Jacobi equations for extremals", associated with control systems and geometric structures. In this way one reduces the problem of construction of the…
This is an expository paper on the theory of gradient flows, and in particular of those PDEs which can be interpreted as gradient flows for the Wasserstein metric on the space of probability measures (a distance induced by optimal…
For free energies of the form \[ F(\mu) = E(\mu) + \sigma\int_\Omega \mu\log\mu\,dx, \quad \sigma > 0, \] we study the Wasserstein gradient flow, a continuity equation also known as mean-field Langevin dynamics, around a stationary state…
We formulate a stochastic equation to model the erosion of a surface with fixed inclination. Because the inclination imposes a preferred direction for material transport, the problem is intrinsically anisotropic. At zeroth order, the…
An open question in the Deep Learning community is why neural networks trained with Gradient Descent generalize well on real datasets even though they are capable of fitting random data. We propose an approach to answering this question…
Backhausz and Szegedy (2019) demonstrated that the almost eigenvectors of random regular graphs converge to Gaussian waves with variance $0\leq \sigma^2\leq 1$. In this paper, we present an alternative proof of this result for the edge…
The study of multivariate extremes is dominated by multivariate regular variation, although it is well known that this approach does not provide adequate distinction between random vectors whose components are not always simultaneously…
In the first part of the paper, comprising section 1 through 6, we introduce a sequence of functions in the tangent bundle TM of any smooth two-dimensional manifold M with smooth Riemannian metric g that correspond to the higher order…
We show that gradient descent converges to a local minimizer, almost surely with random initialization. This is proved by applying the Stable Manifold Theorem from dynamical systems theory.
Gradient descent (GD) on logistic regression has many fascinating properties. When the dataset is linearly separable, it is known that the iterates converge in direction to the maximum-margin separator regardless of how large the step size…
Symmetries are prevalent in deep learning and can significantly influence the learning dynamics of neural networks. In this paper, we examine how exponential symmetries -- a broad subclass of continuous symmetries present in the model…
Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking. These phenomena appear across architectures -- in…
A variant of consensus based distributed gradient descent (\textbf{DGD}) is studied for finite sums of smooth but possibly non-convex functions. In particular, the local gradient term in the fixed step-size iteration of each agent is…