English
Related papers

Related papers: Escaping saddle points without Lipschitz smoothnes…

200 papers

Much is known about when a locally optimal solution depends in a single-valued Lipschitz continuous way on the problem's parameters, including tilt perturbations. Much less is known, however, about when that solution and a uniquely…

Optimization and Control · Mathematics 2024-01-02 Matus Benko , R. Tyrrell Rockafellar

Stochastic gradient descent (SGD) is a prevalent optimization technique for large-scale distributed machine learning. While SGD computation can be efficiently divided between multiple machines, communication typically becomes a bottleneck…

Machine Learning · Computer Science 2021-05-24 Dmitrii Avdiukhin , Grigory Yaroslavtsev

In this paper, we study the problem of escaping from saddle points in smooth nonconvex optimization problems subject to a convex set $\mathcal{C}$. We propose a generic framework that yields convergence to a second-order stationary point of…

Machine Learning · Computer Science 2018-10-10 Aryan Mokhtari , Asuman Ozdaglar , Ali Jadbabaie

Smoothness is crucial for attaining fast rates in first-order optimization. However, many optimization problems in modern machine learning involve non-smooth objectives. Recent studies relax the smoothness assumption by allowing the…

Optimization and Control · Mathematics 2026-02-11 Dingzhi Yu , Wei Jiang , Hongyi Tao , Yuanyu Wan , Lijun Zhang

We give the first polynomial time algorithms for escaping from high-dimensional saddle points under a moderate number of constraints. Given gradient access to a smooth function $f \colon \mathbb R^d \to \mathbb R$ we show that (noisy)…

Machine Learning · Computer Science 2023-04-21 Dmitrii Avdiukhin , Grigory Yaroslavtsev

Adaptive optimization methods (such as Adam) play a major role in LLM pretraining, significantly outperforming Gradient Descent (GD). Recent studies have proposed new smoothness assumptions on the loss function to explain the advantages of…

Machine Learning · Computer Science 2025-12-02 Robin Yadav , Shuo Xie , Tianhao Wang , Zhiyuan Li

$L_0$-smoothness, which has been pivotal to advancing decentralized optimization theory, is often fairly restrictive for modern tasks like deep learning. The recent advent of relaxed $(L_0,L_1)$-smoothness condition enables improved…

Optimization and Control · Mathematics 2025-08-13 Zhanhong Jiang , Aditya Balu , Soumik Sarkar

We introduce a notion of inexact model of a convex objective function, which allows for errors both in the function and in its gradient. For this situation, a gradient method with an adaptive adjustment of some parameters of the model is…

Optimization and Control · Mathematics 2021-10-12 Fedor S. Stonyakin

This paper proposes a family of online second order methods for possibly non-convex stochastic optimizations based on the theory of preconditioned stochastic gradient descent (PSGD), which can be regarded as an enhance stochastic Newton…

Machine Learning · Statistics 2018-05-01 Xi-Lin Li

We present a subgradient method for minimizing non-smooth, non-Lipschitz convex optimization problems. The only structure assumed is that a strictly feasible point is known. We extend the work of Renegar [5] by taking a different…

Optimization and Control · Mathematics 2018-02-28 Benjamin Grimmer

This work considers the non-convex finite sum minimization problem. There are several algorithms for such problems, but existing methods often work poorly when the problem is badly scaled and/or ill-conditioned, and a primary goal of this…

We analyze stochastic gradient algorithms for optimizing nonconvex problems. In particular, our goal is to find local minima (second-order stationary points) instead of just finding first-order stationary points which may be some bad…

Machine Learning · Computer Science 2019-06-24 Zhize Li

In this paper, we study the problem of optimizing a two-layer artificial neural network that best fits a training dataset. We look at this problem in the setting where the number of parameters is greater than the number of sampled points.…

Machine Learning · Computer Science 2017-11-01 Digvijay Boob , Guanghui Lan

We investigate the stochastic optimization problem of minimizing population risk, where the loss defining the risk is assumed to be weakly convex. Compositions of Lipschitz convex functions with smooth maps are the primary examples of such…

Optimization and Control · Mathematics 2018-12-19 Damek Davis , Dmitriy Drusvyatskiy

In part I we considered the problem of convergence to a saddle point of a concave-convex function via gradient dynamics and an exact characterization was given to their asymptotic behaviour. In part II we consider a general class of…

Optimization and Control · Mathematics 2019-08-06 Thomas Holding , Ioannis Lestas

We present an accelerated gradient method for non-convex optimization problems with Lipschitz continuous first and second derivatives. The method requires time $O(\epsilon^{-7/4} \log(1/ \epsilon) )$ to find an $\epsilon$-stationary point,…

Optimization and Control · Mathematics 2017-02-03 Yair Carmon , John C. Duchi , Oliver Hinder , Aaron Sidford

Selecting an effective step-size is a fundamental challenge in first-order optimization, especially for problems with non-Euclidean geometries. This paper presents a novel adaptive step-size strategy for optimization algorithms that rely on…

Optimization and Control · Mathematics 2025-10-14 Abbas Khademi , Antonio Silveti-Falls

An emerging class of trajectory optimization methods enforces collision avoidance by jointly optimizing the robot's configuration and a separating hyperplane. However, as linear separators only apply to convex sets, these methods require…

Robotics · Computer Science 2026-01-15 Shuoye Li , Zhiyuan Song , Yulin Li , Zhihai Bi , Jun Ma

This paper introduces a coordinate descent version of the V\~u-Condat algorithm. By coordinate descent, we mean that only a subset of the coordinates of the primal and dual iterates is updated at each iteration, the other coordinates being…

Optimization and Control · Mathematics 2019-01-17 Olivier Fercoq , Pascal Bianchi

We introduce a general framework for nonlinear stochastic gradient descent (SGD) for the scenarios when gradient noise exhibits heavy tails. The proposed framework subsumes several popular nonlinearity choices, like clipped, normalized,…

Optimization and Control · Mathematics 2022-04-07 Dusan Jakovetic , Dragana Bajovic , Anit Kumar Sahu , Soummya Kar , Nemanja Milosevic , Dusan Stamenkovic