相关论文: Asymptotic behaviour of a family of gradient algor…
This paper investigates asymptotic behaviors of gradient descent algorithms (particularly accelerated gradient descent and stochastic gradient descent) in the context of stochastic optimization arising in statistics and machine learning…
The asymptotic behavior of stochastic gradient algorithms is studied. Relying on results from differential geometry (Lojasiewicz gradient inequality), the single limit-point convergence of the algorithm iterates is demonstrated and…
In a Hilbert space $\mathcal H$, we study the asymptotic behaviour, as time variable $t$ goes to $+\infty$, of nonautonomous gradient-like dynamical systems involving inertia and multiscale features. Given $\mathcal H$ a general Hilbert…
The asymptotic behavior of the stochastic gradient algorithm with a biased gradient estimator is analyzed. Relying on arguments based on the dynamic system theory (chain-recurrence) and the differential geometry (Yomdin theorem and…
In this paper, we investigate a general class of stochastic gradient descent (SGD) algorithms, called Conditioned SGD, based on a preconditioning of the gradient direction. Using a discrete-time approach with martingale tools, we establish…
Considering the constrained stochastic optimization problem over a time-varying random network, where the agents are to collectively minimize a sum of objective functions subject to a common constraint set, we investigate asymptotic…
We study the convergence of a random iterative sequence of a family of operators on infinite dimensional Hilbert spaces, inspired by the Stochastic Gradient Descent (SGD) algorithm in the case of the noiseless regression, as studied in [1].…
We consider the asymptotic behavior of a family of gradient methods, which include the steepest descent and minimal gradient methods as special instances. It is proved that each method in the family will asymptotically zigzag between two…
We develop a new asymptotic method for the analysis of matrix Riemann-Hilbert problems. Our method is a generalization of the steepest descent method first proposed by Deift and Zhou; however our method systematically handles jump matrices…
Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-studied when…
This paper studies some asymptotic properties of adaptive algorithms widely used in optimization and machine learning, and among them Adagrad and Rmsprop, which are involved in most of the blackbox deep learning algorithms. Our setup is the…
We study a nonlinear semigroup associated to a nonexpansive mapping on a Hadamard space and establish its weak convergence to a fixed point. A discrete-time counterpart of such a semigroup, the proximal point algorithm, turns out to have…
Stochastic gradient descent (SGD) has been studied extensively over the past decades due to its simplicity and broad applicability in machine learning. In this work, we analyze the local behavior of gradient descent and stochastic gradient…
We study the asymptotic behavior of second-order algorithms mixing Newton's method and inertial gradient descent in non-convex landscapes. We show that, despite the Newtonian behavior of these methods, they almost always escape strict…
We establish formulae for the asymptotic growth (with respect to the scaling dimension) of the number of operators in effective field theory, or equivalently the number of $S$-matrix elements, in arbitrary spacetime dimensions and with…
We give an overview of operator-theoretic tools that have recently proved useful in the analysis of boundary-value and transmission problems for second-order partial differential equations, with a view to addressing, in particular, the…
We study the asymptotic behavior of the trajectory of a nonautonomous evolution equation governed by a quasi-nonexpansive operator in Hilbert spaces. We prove the weak convergence of the trajectory to a fixed point of the operator by…
Adaptive gradient methods, such as AdaGrad, have become fundamental tools in deep learning. Despite their widespread use, the asymptotic convergence of AdaGrad remains poorly understood in non-convex scenarios. In this work, we present the…
We characterise asymptotic behaviour of families of symmetric orthonormal polynomials whose recursion coefficients satisfy certain conditions, satisfied for example by the (normalised) Hermite polynomials. More generally, these conditions…
This paper proposes a two-point inertial proximal point algorithm to find zero of maximal monotone operators in Hilbert spaces. We obtain weak convergence results and non-asymptotic $O(1/n)$ convergence rate of our proposed algorithm in…