English
Related papers

Related papers: Reachability of gradient descent

200 papers

An approximation, in the sense of $\Gamma$-convergence and in any dimension $d\geq1$, of Griffith-type functionals, with $p-$growth ($p>1$) in the symmetrized gradient, is provided by means of a sequence of non-local integral functionals…

Analysis of PDEs · Mathematics 2021-02-05 Giovanni Scilla , Francesco Solombrino

The convergence of stochastic gradient descent is highly dependent on the step-size, especially on non-convex problems such as neural network training. Step decay step-size schedules (constant and then cut) are widely used in practice…

Optimization and Control · Mathematics 2021-02-19 Xiaoyu Wang , Sindri Magnússon , Mikael Johansson

Gradient descent is a simple and widely used optimization method for machine learning. For homogeneous linear classifiers applied to separable data, gradient descent has been shown to converge to the maximal margin (or equivalently, the…

Machine Learning · Statistics 2019-07-30 Denali Molitor , Deanna Needell , Rachel Ward

We consider the problem of unconstrained minimization of finite sums of functions. We propose a simple, yet, practical way to incorporate variance reduction techniques into SignSGD, guaranteeing convergence that is similar to the full sign…

Optimization and Control · Mathematics 2023-05-23 Evgenii Chzhen , Sholom Schechtman

We show that in a variety of large-scale deep learning scenarios the gradient dynamically converges to a very small subspace after a short period of training. The subspace is spanned by a few top eigenvectors of the Hessian (equal to the…

Machine Learning · Computer Science 2018-12-13 Guy Gur-Ari , Daniel A. Roberts , Ethan Dyer

We propose a mini-batching scheme for improving the theoretical complexity and practical performance of semi-stochastic gradient descent applied to the problem of minimizing a strongly convex composite function represented as the sum of an…

Machine Learning · Computer Science 2014-10-20 Jakub Konečný , Jie Liu , Peter Richtárik , Martin Takáč

This paper presents an extension of stochastic gradient descent for the minimization of Lipschitz continuous loss functions. Our motivation is for use in non-smooth non-convex stochastic optimization problems, which are frequently…

Optimization and Control · Mathematics 2022-10-05 Michael R. Metel , Akiko Takeda

We introduce a perturbed preconditioned gradient descent (PPGD) method for the unconstrained minimization of a strongly convex objective $G$ with a locally Lipschitz continuous gradient. We assume that $G(v)=E(v)+F(v)$ and that the gradient…

Optimization and Control · Mathematics 2025-12-23 Jea-Hyun Park , Abner J. Salgado , Steven M. Wise

In this paper, we propose a simple, fast and easy to implement algorithm LOSSGRAD (locally optimal step-size in gradient descent), which automatically modifies the step-size in gradient descent during neural networks training. Given a…

Machine Learning · Computer Science 2019-11-26 Bartosz Wójcik , Łukasz Maziarka , Jacek Tabor

We analyze the convergence of a nonlocal gradient descent method for minimizing a class of high-dimensional non-convex functions, where a directional Gaussian smoothing (DGS) is proposed to define the nonlocal gradient (also referred to as…

Optimization and Control · Mathematics 2023-02-14 Hoang Tran , Qiang Du , Guannan Zhang

Goldstein's 1977 idealized iteration for minimizing a Lipschitz objective fixes a distance - the step size - and relies on a certain approximate subgradient. That "Goldstein subgradient" is the shortest convex combination of objective…

Optimization and Control · Mathematics 2024-05-22 Siyu Kong , Adrian S. Lewis

In this paper, we consider two variants of the concept of sharp minimum for mathematical programming problems with quasiconvex objective function and inequality constraints. It investigated the problem of describing a variant of a simple…

Optimization and Control · Mathematics 2023-12-29 S. M. Puchinin , E. R. Korolkov , F. S. Stonyakin , M. S. Alkousa , A. A Vyguzov

We provide larger step-size restrictions for which gradient descent based algorithms (almost surely) avoid strict saddle points. In particular, consider a twice differentiable (non-convex) objective function whose gradient has Lipschitz…

Machine Learning · Statistics 2019-08-06 Hayden Schaeffer , Scott G. McCalla

Gradient descent (GD) on logistic regression has many fascinating properties. When the dataset is linearly separable, it is known that the iterates converge in direction to the maximum-margin separator regardless of how large the step size…

Machine Learning · Computer Science 2025-07-16 Si Yi Meng , Baptiste Goujaud , Antonio Orvieto , Christopher De Sa

We show that the vanishing stepsize subgradient method -- widely adopted for machine learning applications -- can display rather messy behavior even in the presence of favorable assumptions. We establish that convergence of bounded…

Optimization and Control · Mathematics 2020-07-24 Rodolfo Rios-Zertuche

In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event…

Machine Learning · Computer Science 2024-11-20 Steffen Dereich , Sebastian Kassing

In this work, we analyze the global convergence property of coordinate gradient descent with random choice of coordinates and stepsizes for non-convex optimization problems. Under generic assumptions, we prove that the algorithm iterate…

Optimization and Control · Mathematics 2022-12-01 Ziang Chen , Yingzhou Li , Jianfeng Lu

In this paper some adaptive mirror descent algorithms for problems of minimization convex objective functional with several convex Lipschitz (generally, non-smooth) functional constraints are considered. It is shown that the methods are…

Optimization and Control · Mathematics 2018-12-20 F. S. Stonyakin , M . S. Alkousa , A. A. Titov

We show that adaptive proximal gradient methods for convex problems are not restricted to traditional Lipschitzian assumptions. Our analysis reveals that a class of linesearch-free methods is still convergent under mere local H\"older…

Optimization and Control · Mathematics 2024-07-08 Konstantinos A. Oikonomidis , Emanuel Laude , Puya Latafat , Andreas Themelis , Panagiotis Patrinos

We consider a stochastic version of the proximal point algorithm for optimization problems posed on a Hilbert space. A typical application of this is supervised learning. While the method is not new, it has not been extensively analyzed in…

Optimization and Control · Mathematics 2021-09-28 Monika Eisenmann , Tony Stillfjord , Måns Williamson