Related papers: One-dimensional System Arising in Stochastic Gradi…
We present a scheme improving the minimum-mode following method for finding first order saddle points by confining the displacements of atoms to the subset of those subject to the largest force. By doing so it is ensured that the…
A central challenge to many fields of science and engineering involves minimizing non-convex error functions over continuous, high dimensional spaces. Gradient descent or quasi-Newton methods are almost ubiquitously used to perform such…
Gaussian processes are a powerful framework for quantifying uncertainty and for sequential decision-making but are limited by the requirement of solving linear systems. In general, this has a cubic cost in dataset size and is sensitive to…
We prove existence and pathwise uniqueness results for four different types of stochastic differential equations (SDEs) perturbed by the past maximum process and/or the local time at zero. Along the first three studies, the coefficients are…
In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of…
This paper presents the first sufficient conditions that guarantee the stability and almost sure convergence of multi-timescale stochastic approximation (SA) iterates. It extends the existing results on one-timescale and two-timescale SA…
In this paper, we focus on solving a class of constrained non-convex non-concave saddle point problems in a decentralized manner by a group of nodes in a network. Specifically, we assume that each node has access to a summand of a global…
In this paper, we adapt proximal incremental aggregated gradient methods to saddle point problems, which is motivated by decoupling linear transformations in regularized empirical risk minimization models. First, the Primal-Dual Proximal…
Satisfaction of the strict saddle property has become a standard assumption in non-convex optimization, and it ensures that many first-order optimization algorithms will almost always escape saddle points. However, functions exist in…
We study the learning dynamics of a multi-pass, mini-batch Stochastic Gradient Descent (SGD) procedure for empirical risk minimization in high-dimensional multi-index models with isotropic random data. In an asymptotic regime where the…
In this paper, we study degenerate entry-exit problems associated with planar slow-fast systems having an invariant line $\{(x,y)\,:\,y=0\}$ with a turning point at $x=0$. The degeneracy stems from the fact that the slow flow has a…
Stochastic gradient descent (SGD) is a standard optimization method to minimize a training error with respect to network parameters in modern neural network learning. However, it typically suffers from proliferation of saddle points in the…
We establish Freidlin-Wentzell results for a nonlinear ordinary differential equation starting close to the stable state $0$, say, subject to a perturbation by a stochastic integral which is driven by an $\varepsilon$-small and…
Subspace learning and matrix factorization problems have great many applications in science and engineering, and efficient algorithms are critical as dataset sizes continue to grow. Many relevant problem formulations are non-convex, and in…
This paper presents an algorithm for the efficient approximation of the saddle-extremum persistence diagram of a scalar field. Vidal et al. introduced recently a fast algorithm for such an approximation (by interrupting a progressive…
In this paper, we study solutions $u$ of parabolic systems in divergence form with zero Dirichlet boundary conditions in the upper-half cylinder $Q_1^+\subset \mathbb{R}^{n+1}$, where the coefficients are weighted by $x_n^\alpha$,…
Here we present a multiscale method to calculate the saddle point associated with the effective dynamics arising from a stochastic system which couples slow deterministic drift and fast stochastic dynamics. This problem is motivated by the…
Let $B$ be a $d$-dimensional Gaussian process on $\mathbb{R}$, where the component are independents copies of a scalar Gaussian process $B_0$ on $\mathbb{R}_+$ with a given general variance function…
We consider the problem of convergence to a saddle point of a concave-convex function via gradient dynamics. Since first introduced by Arrow, Hurwicz and Uzawa in [1] such dynamics have been extensively used in diverse areas, there are,…
The purpose of this paper is to study some properties of solutions to one dimensional as well as multidimensional stochastic differential equations (SDEs in short) with super-linear growth conditions on the coefficients. Taking inspiration…