English
Related papers

Related papers: Functional Central Limit Theorem for Stochastic Gr…

200 papers

We propose a new design strategy for extremum seeking control for a multi-dimensional single-integrator system in the presence of local extrema. The proposed method employs suitably designed sinusoidal dither signals, which force the…

Optimization and Control · Mathematics 2026-03-03 Raik Suttner , Christian Ebenbauer , Sergey Dashkovskiy

We study the asymptotic behavior of stochastic hyperbolic parabolic equations with slow and fast time scales. Both the strong and weak convergence in the averaging principe are established, which can be viewed as a functional law of large…

Probability · Mathematics 2020-11-12 Michael Röckner , Longjie Xie , Li Yang

This paper is concerned with convergence of stochastic gradient algorithms with momentum terms in the nonconvex setting. A class of stochastic momentum methods, including stochastic gradient descent, heavy ball, and Nesterov's accelerated…

Optimization and Control · Mathematics 2021-10-01 Zixuan Wang , Shanjian Tang

The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-scale data. In this context, momentum-based stochastic optimisation algorithms have become…

Optimization and Control · Mathematics 2024-11-06 Kexin Jin , Jonas Latz , Chenguang Liu , Alessandro Scagliotti

We consider the adjacency matrix $A$ of a large random graph and study fluctuations of the function $f_n(z,u)=\frac{1}{n}\sum_{k=1}^n\exp\{-uG_{kk}(z)\}$ with $G(z)=(z-iA)^{-1}$. We prove that the moments of fluctuations normalized by…

Mathematical Physics · Physics 2015-05-14 M. Shcherbina , B. Tirozzi

Stochastic gradient descent (SGD) is a popular algorithm for minimizing objective functions that arise in machine learning. For constant step-sized SGD, the iterates form a Markov chain on a general state space. Focusing on a class of…

Optimization and Control · Mathematics 2025-03-26 David Shirokoff , Philip Zaleski

In machine learning, stochastic gradient descent (SGD) is widely deployed to train models using highly non-convex objectives with equally complex noise models. Unfortunately, SGD theory often makes restrictive assumptions that fail to…

Machine Learning · Computer Science 2022-10-11 Vivak Patel , Shushu Zhang , Bowen Tian

We establish central and non-central limit theorems for sequences of functionals of the Gaussian output of an infinitely-wide random neural network on the d-dimensional sphere . We show that the asymptotic behaviour of these functionals as…

Probability · Mathematics 2026-04-24 Simmaco Di Lillo , Leonardo Maini , Domenico Marinucci

We consider the minimization of non-convex quadratic forms regularized by a cubic term, which exhibit multiple saddle points and poor local minima. Nonetheless, we prove that, under mild assumptions, gradient descent approximates the…

Optimization and Control · Mathematics 2022-08-31 Yair Carmon , John C. Duchi

Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch methods tend to converge to sharp minimizers has received…

Machine Learning · Statistics 2018-12-04 Xiaowu Dai , Yuhua Zhu

We propose to optimize neural networks with a uniformly-distributed random learning rate. The associated stochastic gradient descent algorithm can be approximated by continuous stochastic equations and analyzed within the Fokker-Planck…

Machine Learning · Computer Science 2020-10-13 Daniele Musso

We study the distributed stochastic compositional optimization problems over directed communication networks in which agents privately own a stochastic compositional objective function and collaborate to minimize the sum of all objective…

Optimization and Control · Mathematics 2022-03-22 Shengchao Zhao , Yongchao Liu

This paper is devoted to the non-asymptotic control of the mean-squared error for the Ruppert-Polyak stochastic averaged gradient descent introduced in the seminal contributions of [Rup88] and [PJ92]. In our main results, we establish…

Statistics Theory · Mathematics 2017-09-12 Sébastien Gadat , Fabien Panloup

Decentralized optimization is a powerful paradigm that finds applications in engineering and learning design. This work studies decentralized composite optimization problems with non-smooth regularization terms. Most existing gradient-based…

Optimization and Control · Mathematics 2019-10-29 Sulaiman A. Alghunaim , Kun Yuan , Ali H. Sayed

We show that gradient descent converges to a local minimizer, almost surely with random initialization. This is proved by applying the Stable Manifold Theorem from dynamical systems theory.

Machine Learning · Statistics 2016-03-07 Jason D. Lee , Max Simchowitz , Michael I. Jordan , Benjamin Recht

Limit distributions for the greatest convex minorant and its derivative are considered for a general class of stochastic processes including partial sum processes and empirical processes, for independent, weakly dependent and long range…

Statistics Theory · Mathematics 2016-08-16 D. Anevski , O. Hössjer

We consider a recursive algorithm to construct an aggregated estimator from a finite number of base decision rules in the classification problem. The estimator approximately minimizes a convex risk functional under the l1-constraint. It is…

Statistics Theory · Mathematics 2007-06-13 Anatoli Juditsky , Alexander Nazin , Alexandre Tsybakov , Nicolas Vayatis

In this paper we present a novel randomized block coordinate descent method for the minimization of a convex composite objective function. The method uses (approximate) partial second-order (curvature) information, so that the algorithm…

Optimization and Control · Mathematics 2015-05-11 Kimon Fountoulakis , Rachael Tappenden

Consider d uniformly random permutation matrices on n labels. Consider the sum of these matrices along with their transposes. The total can be interpreted as the adjacency matrix of a random regular graph of degree 2d on n vertices. We…

Probability · Mathematics 2019-09-25 Ioana Dumitriu , Tobias Johnson , Soumik Pal , Elliot Paquette

We study the statistical and computational complexities of the Polyak step size gradient descent algorithm under generalized smoothness and Lojasiewicz conditions of the population loss function, namely, the limit of the empirical loss…

Machine Learning · Computer Science 2021-10-18 Tongzheng Ren , Fuheng Cui , Alexia Atsidakou , Sujay Sanghavi , Nhat Ho
‹ Prev 1 8 9 10 Next ›