English
Related papers

Related papers: On the diameter of subgradient sequences in o-mini…

200 papers

We consider a stochastic version of the proximal point algorithm for optimization problems posed on a Hilbert space. A typical application of this is supervised learning. While the method is not new, it has not been extensively analyzed in…

Optimization and Control · Mathematics 2021-09-28 Monika Eisenmann , Tony Stillfjord , Måns Williamson

If we consider a sequence of warped product length spaces, what conditions on the sequence of warping functions implies compactness of the sequence of distance functions? In particular, we want to know when a subsequence converges to a well…

Differential Geometry · Mathematics 2024-09-12 Brian Allen , Bryan Sanchez , Yahaira Torres

We analyze the stochastic proximal subgradient descent in the case where the objective functions are path differentiable and verify a Sard-type condition. While the accumulation set may not be reduced to unique point, we show that the time…

Optimization and Control · Mathematics 2022-05-25 Sholom Schechtman

Stochastic Gradient Descent (SGD) plays a central role in modern machine learning. While there is extensive work on providing error upper bound for SGD, not much is known about SGD error lower bound. In this paper, we study the convergence…

Optimization and Control · Mathematics 2019-10-21 Zhiyan Ding , Yiding Chen , Qin Li , Xiaojin Zhu

The convergence theory for the gradient sampling algorithm is extended to directionally Lipschitz functions. Although directionally Lipschitz functions are not necessarily locally Lipschitz, they are almost everywhere differentiable and…

Optimization and Control · Mathematics 2021-07-13 James V. Burke , Qiuying Lin

In this work, we establish regularity results for minimizers of the energy functional associated with the thin obstacle problem in Orlicz spaces. More precisely, we prove the Lipschitz continuity and the H\"older continuity of the gradient…

Analysis of PDEs · Mathematics 2026-02-05 Junior da Silva Bessa , Paulo Henryque da Costa Silva , Alan Pio Sousa

In this paper we study local error bound moduli for a locally Lipschitz and regular function via its outer limiting subdifferential set. We show that the distance of 0 from the outer limiting subdifferential of the support function of the…

Optimization and Control · Mathematics 2016-08-12 Minghua Li , Kaiwen Meng , Xiaoqi Yang

It is well-known that the convergence of a family of smooth functions does not imply the convergence of its gradients. In this work, we show that if the family is definable in an o-minimal structure (for instance semialgebraic, subanalytic,…

Optimization and Control · Mathematics 2026-02-17 Sholom Schechtman

We propose a single time-scale stochastic subgradient method for constrained optimization of a composition of several nonsmooth and nonconvex functions. The functions are assumed to be locally Lipschitz and differentiable in a generalized…

Optimization and Control · Mathematics 2020-12-22 Andrzej Ruszczynski

Using a geometric argument, we show that under a reasonable continuity condition, the Clarke subdifferential of a semi-algebraic (or more generally stratifiable) directionally Lipschitzian function admits a simple form: the normal cone to…

Optimization and Control · Mathematics 2012-11-16 Dmitriy Drusvyatskiy , Alexander D. Ioffe , Adrian S. Lewis

We construct a Lipschitz truncation which approximates functions of bounded variation in the area-strict metric. The Lipschitz truncation changes the original function only on a small set similar to Lusin's theorem. Previous results could…

Analysis of PDEs · Mathematics 2019-08-29 Dominic Breit , Lars Diening , Franz Gmeineder

This work considers gradient descent for L-smooth convex optimization with stepsizes larger than the classic regime where descent can be ensured. The stepsize schedules considered are similar to but differ slightly from the recent silver…

Optimization and Control · Mathematics 2024-04-15 Benjamin Grimmer , Kevin Shu , Alex L. Wang

We demonstrate that for strongly log-convex densities whose potentials are discontinuous on manifolds, the ULA algorithm converges with stepsize bias of order $1/2$ in Wasserstein-p distance. Our resulting bound is then of the same order as…

Probability · Mathematics 2023-12-05 Tim Johnston , Sotirios Sabanis

Motivated by conforming finite element methods for elliptic problems of second order, we analyze the approximation of the gradient of a target function by continuous piecewise polynomial functions over a simplicial mesh. The main result is…

Numerical Analysis · Mathematics 2018-03-07 Andreas Veeser

It is well known that both gradient descent and stochastic coordinate descent achieve a global convergence rate of $O(1/k)$ in the objective value, when applied to a scheme for minimizing a Lipschitz-continuously differentiable,…

Optimization and Control · Mathematics 2019-05-15 Ching-pei Lee , Stephen J. Wright

Consider the problem of minimizing functions that are Lipschitz and strongly convex, but not necessarily differentiable. We prove that after $T$ steps of stochastic gradient descent, the error of the final iterate is $O(\log(T)/T)$ with…

Machine Learning · Computer Science 2018-12-14 Nicholas J. A. Harvey , Christopher Liaw , Yaniv Plan , Sikander Randhawa

This paper introduces a novel approach to enhance the performance of the stochastic gradient descent (SGD) algorithm by incorporating a modified decay step size based on $\frac{1}{\sqrt{t}}$. The proposed step size integrates a logarithmic…

Machine Learning · Computer Science 2023-09-06 M. Soheil Shamaee , S. Fathi Hafshejani

Goldstein's 1977 idealized iteration for minimizing a Lipschitz objective fixes a distance - the step size - and relies on a certain approximate subgradient. That "Goldstein subgradient" is the shortest convex combination of objective…

Optimization and Control · Mathematics 2024-05-22 Siyu Kong , Adrian S. Lewis

The convergence of stochastic gradient descent is highly dependent on the step-size, especially on non-convex problems such as neural network training. Step decay step-size schedules (constant and then cut) are widely used in practice…

Optimization and Control · Mathematics 2021-02-19 Xiaoyu Wang , Sindri Magnússon , Mikael Johansson

The mini-batch stochastic gradient descent (SGD) algorithm is widely used in training machine learning models, in particular deep learning models. We study SGD dynamics under linear regression and two-layer linear networks, with an easy…

Optimization and Control · Mathematics 2020-04-29 Xin Qian , Diego Klabjan