English
Related papers

Related papers: Central limit theorems for stochastic gradient des…

200 papers

A class of random recursive sequences (Y_n) with slowly varying variances as arising for parameters of random trees or recursive algorithms leads after normalizations to degenerate limit equations of the form X\stackrel{L}{=}X. For…

Probability · Mathematics 2016-09-07 Ralph Neininger , Ludger Ruschendorf

This paper enhances the result of the work [G. Kozma, B. T\'oth, Ann. Probab. vol. 45 (2017) 4307-4347] . We prove the central limit theorem (in probability w.r.t. the environment) for the displacement of a random walker in divergence-free…

Probability · Mathematics 2026-02-19 Bálint Tóth

General Central limit theorem deals with weak limits (in type) of sums of row-elements of array random variables. In some situations as in the invariance principle problem, the sums may include only parts of the row-elements. For strictly…

This work investigates stepsize-based acceleration of gradient descent with {\em anytime} convergence guarantees. For smooth (non-strongly) convex optimization, we propose a stepsize schedule that allows gradient descent to achieve…

Machine Learning · Computer Science 2024-12-10 Zihan Zhang , Jason D. Lee , Simon S. Du , Yuxin Chen

The question of whether the central limit theorem (CLT) holds for the total number of edges in exponential random graph models (ERGMs) in the subcritical region of parameters has remained an open problem. In this paper, we establish the…

Probability · Mathematics 2025-04-09 Xiao Fang , Song-Hao Liu , Qi-Man Shao , Yi-Kun Zhao

The distribution of descents in fixed conjugacy classes of $S_n$ has been studied, and it is shown that its moments have interesting properties. Fulman proved that the descent numbers of permutations in conjugacy classes with large cycles…

Combinatorics · Mathematics 2018-03-29 Gene B. Kim , Sangchul Lee

The general model of coagulation is considered. For basic classes of unbounded coagulation kernels the central limit theorem (CLT) is obtained for the fluctuations around the dynamic law of large numbers (LLN). A rather precise rate of…

Probability · Mathematics 2022-05-03 Vassili Kolokoltsov

We consider optimizing a function smooth convex function $f$ that is the average of a set of differentiable functions $f_i$, under the assumption considered by Solodov [1998] and Tseng [1998] that the norm of each gradient $f_i'$ is bounded…

Optimization and Control · Mathematics 2013-08-30 Mark Schmidt , Nicolas Le Roux

We study the iteration complexity of stochastic gradient descent (SGD) for minimizing the gradient norm of smooth, possibly nonconvex functions. We provide several results, implying that the $\mathcal{O}(\epsilon^{-4})$ upper bound of…

Machine Learning · Computer Science 2021-07-30 Yoel Drori , Ohad Shamir

A vast literature on convergence guarantees for gradient descent and derived methods exists at the moment. However, a simple practical situation remains unexplored: when a fixed step size is used, can we expect gradient descent to converge…

Machine Learning · Computer Science 2024-12-10 Alexandru Crăciun , Debarghya Ghoshdastidar

In machine learning, stochastic gradient descent (SGD) is widely deployed to train models using highly non-convex objectives with equally complex noise models. Unfortunately, SGD theory often makes restrictive assumptions that fail to…

Machine Learning · Computer Science 2022-10-11 Vivak Patel , Shushu Zhang , Bowen Tian

In this paper, we aim to study the asymptotic behavior for multi-scale McKean-Vlasov stochastic dynamical systems. Firstly, we obtain a central limit type theorem, i.e, the deviation between the slow component $X^{\varepsilon}$ and the…

Probability · Mathematics 2023-06-02 Wei Hong , Shihu Li , Wei Liu , Xiaobin Sun

In this paper, we consider a general stochastic optimization problem which is often at the core of supervised learning, such as deep learning and linear classification. We consider a standard stochastic gradient descent (SGD) method with a…

Machine Learning · Statistics 2018-12-27 Lam M. Nguyen , Nam H. Nguyen , Dzung T. Phan , Jayant R. Kalagnanam , Katya Scheinberg

Suppose $B_i:= B(p,r_i)$ are nested balls of radius $r_i$ about a point $p$ in a dynamical system $(T,X,\mu)$. The question of whether $T^i x\in B_i$ infinitely often (i. o.) for $\mu$ a.e.\ $x$ is often called the shrinking target problem.…

Dynamical Systems · Mathematics 2015-06-16 Nicolai Haydn , Matthew Nicol , Sandro Vaienti , Licheng Zhang

Stochastic optimization via Stochastic Gradient Descent (SGD) is a fundamental problem in statistics and optimization. This paper revisits Stochastic Gradient Descent (SGD) for strongly convex objectives, establishing tight, uniform-in-time…

Optimization and Control · Mathematics 2026-03-19 Kang Chen , Yasong Feng , Tianyu Wang

We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…

Machine Learning · Computer Science 2025-01-15 Aaron Mishkin , Ahmed Khaled , Yuanhao Wang , Aaron Defazio , Robert M. Gower

Well-posedness of a reversible variant of the Gray-Scott model is shown, along with the convergence of each trajectory to one of the two spatially homogeneous steady states. The principle of linearized stability provides the local…

Analysis of PDEs · Mathematics 2025-12-04 Philippe Laurençot , Christoph Walker

A line of recent works established that when training linear predictors over separable data, using gradient methods and exponentially-tailed losses, the predictors asymptotically converge in direction to the max-margin predictor. As a…

Machine Learning · Computer Science 2020-09-11 Ohad Shamir

We give optimal convergence rates in the central limit theorem for a large class of martingale difference sequences with bounded third moments. The rates depend on the behaviour of the conditional variances and for stationary sequences the…

Probability · Mathematics 2007-05-23 Mohamed El Machkouri , Lahcen Ouchti

We show that accelerated gradient descent, averaged gradient descent and the heavy-ball method for non-strongly-convex problems may be reformulated as constant parameter second-order difference equation algorithms, where stability of the…

Machine Learning · Statistics 2015-04-08 Nicolas Flammarion , Francis Bach
‹ Prev 1 4 5 6 7 8 10 Next ›