Related papers: Central limit theorems for stochastic gradient des…
A class of random recursive sequences (Y_n) with slowly varying variances as arising for parameters of random trees or recursive algorithms leads after normalizations to degenerate limit equations of the form X\stackrel{L}{=}X. For…
This paper enhances the result of the work [G. Kozma, B. T\'oth, Ann. Probab. vol. 45 (2017) 4307-4347] . We prove the central limit theorem (in probability w.r.t. the environment) for the displacement of a random walker in divergence-free…
General Central limit theorem deals with weak limits (in type) of sums of row-elements of array random variables. In some situations as in the invariance principle problem, the sums may include only parts of the row-elements. For strictly…
This work investigates stepsize-based acceleration of gradient descent with {\em anytime} convergence guarantees. For smooth (non-strongly) convex optimization, we propose a stepsize schedule that allows gradient descent to achieve…
The question of whether the central limit theorem (CLT) holds for the total number of edges in exponential random graph models (ERGMs) in the subcritical region of parameters has remained an open problem. In this paper, we establish the…
The distribution of descents in fixed conjugacy classes of $S_n$ has been studied, and it is shown that its moments have interesting properties. Fulman proved that the descent numbers of permutations in conjugacy classes with large cycles…
The general model of coagulation is considered. For basic classes of unbounded coagulation kernels the central limit theorem (CLT) is obtained for the fluctuations around the dynamic law of large numbers (LLN). A rather precise rate of…
We consider optimizing a function smooth convex function $f$ that is the average of a set of differentiable functions $f_i$, under the assumption considered by Solodov [1998] and Tseng [1998] that the norm of each gradient $f_i'$ is bounded…
We study the iteration complexity of stochastic gradient descent (SGD) for minimizing the gradient norm of smooth, possibly nonconvex functions. We provide several results, implying that the $\mathcal{O}(\epsilon^{-4})$ upper bound of…
A vast literature on convergence guarantees for gradient descent and derived methods exists at the moment. However, a simple practical situation remains unexplored: when a fixed step size is used, can we expect gradient descent to converge…
In machine learning, stochastic gradient descent (SGD) is widely deployed to train models using highly non-convex objectives with equally complex noise models. Unfortunately, SGD theory often makes restrictive assumptions that fail to…
In this paper, we aim to study the asymptotic behavior for multi-scale McKean-Vlasov stochastic dynamical systems. Firstly, we obtain a central limit type theorem, i.e, the deviation between the slow component $X^{\varepsilon}$ and the…
In this paper, we consider a general stochastic optimization problem which is often at the core of supervised learning, such as deep learning and linear classification. We consider a standard stochastic gradient descent (SGD) method with a…
Suppose $B_i:= B(p,r_i)$ are nested balls of radius $r_i$ about a point $p$ in a dynamical system $(T,X,\mu)$. The question of whether $T^i x\in B_i$ infinitely often (i. o.) for $\mu$ a.e.\ $x$ is often called the shrinking target problem.…
Stochastic optimization via Stochastic Gradient Descent (SGD) is a fundamental problem in statistics and optimization. This paper revisits Stochastic Gradient Descent (SGD) for strongly convex objectives, establishing tight, uniform-in-time…
We develop new sub-optimality bounds for gradient descent (GD) that depend on the conditioning of the objective along the path of optimization rather than on global, worst-case constants. Key to our proofs is directional smoothness, a…
Well-posedness of a reversible variant of the Gray-Scott model is shown, along with the convergence of each trajectory to one of the two spatially homogeneous steady states. The principle of linearized stability provides the local…
A line of recent works established that when training linear predictors over separable data, using gradient methods and exponentially-tailed losses, the predictors asymptotically converge in direction to the max-margin predictor. As a…
We give optimal convergence rates in the central limit theorem for a large class of martingale difference sequences with bounded third moments. The rates depend on the behaviour of the conditional variances and for stationary sequences the…
We show that accelerated gradient descent, averaged gradient descent and the heavy-ball method for non-strongly-convex problems may be reformulated as constant parameter second-order difference equation algorithms, where stability of the…