English
Related papers

Related papers: Can Single-Shuffle SGD be Better than Reshuffling …

200 papers

Large-scale constrained optimization problems are at the core of many tasks in control, signal processing, and machine learning. Notably, problems with functional constraints arise when, beyond a performance{\nobreakdash-}centric goal…

Optimization and Control · Mathematics 2025-05-15 Antesh Upadhyay , Sang Bin Moon , Abolfazl Hashemi

In the realm of big data and machine learning, data-parallel, distributed stochastic algorithms have drawn significant attention in the present days.~While the synchronous versions of these algorithms are well understood in terms of their…

Optimization and Control · Mathematics 2020-04-07 Atal Narayan Sahu , Aritra Dutta , Aashutosh Tiwari , Peter Richtárik

Based on SGD, previous works have proposed many algorithms that have improved convergence speed and generalization in stochastic optimization, such as SGDm, AdaGrad, Adam, etc. However, their convergence analysis under non-convex conditions…

Machine Learning · Computer Science 2024-02-05 Yichuan Deng , Zhao Song , Chiwun Yang

Consider solving large sparse range symmetric singular linear systems $ A {\bf x}= {\bf b} $ which arise, for instance, in the discretization of convection diffusion equations with periodic boundary conditions, and partial differential…

Numerical Analysis · Mathematics 2022-11-02 Kota Sugihara , Ken Hayami , Liao Zeyu

A Gilbert-Shannon-Reeds (GSR) shuffle is performed on a deck of $N$ cards by cutting the top $n\sim Bin(N,1/2)$ cards and interleaving the two resulting piles uniformly at random. The celebrated "Seven shuffles suffice" theorem of…

Probability · Mathematics 2025-10-28 Mark Sellke , Jialu Shi , Jiamin Wang

This paper proposes a thorough theoretical analysis of Stochastic Gradient Descent (SGD) with non-increasing step sizes. First, we show that the recursion defining SGD can be provably approximated by solutions of a time inhomogeneous…

Optimization and Control · Mathematics 2021-02-02 Xavier Fontaine , Valentin De Bortoli , Alain Durmus

We study the mixing properties for stochastic accelerated gradient descent (SAGD) on least-squares regression. First, we show that stochastic gradient descent (SGD) and SAGD are simulating the same invariant distribution. Motivated by this,…

Optimization and Control · Mathematics 2019-11-01 Peiyuan Zhang , Hadi Daneshmand , Thomas Hofmann

We study the problem of solving strongly convex and smooth unconstrained optimization problems using stochastic first-order algorithms. We devise a novel algorithm, referred to as Recursive One-Over-T SGD (ROOT-SGD), based on an easily…

Optimization and Control · Mathematics 2024-09-19 Chris Junchi Li , Wenlong Mou , Martin J. Wainwright , Michael I. Jordan

We consider symmetric and Hermitian random matrices whose entries are independent and symmetric random variables with an arbitrary variance pattern. Under a novel Short-to-Long Mixing condition, which is sharp in the sense that it precludes…

Probability · Mathematics 2025-11-12 Dang-Zheng Liu , Guangyi Zou

Stochastic Gradient Descent (SGD) is one of the simplest and most popular stochastic optimization methods. While it has already been theoretically studied for decades, the classical analysis usually required non-trivial smoothness…

Machine Learning · Computer Science 2013-01-01 Ohad Shamir , Tong Zhang

Stochastic gradient descent (SGD) on a low-rank factorization is commonly employed to speed up matrix problems including matrix completion, subspace tracking, and SDP relaxation. In this paper, we exhibit a step size scheme for SGD on a…

Machine Learning · Computer Science 2015-02-11 Christopher De Sa , Kunle Olukotun , Christopher Ré

In this paper, we propose a simple variant of the original SVRG, called variance reduced stochastic gradient descent (VR-SGD). Unlike the choices of snapshot and starting points in SVRG and its proximal variant, Prox-SVRG, the two vectors…

Machine Learning · Computer Science 2018-10-31 Fanhua Shang , Kaiwen Zhou , Hongying Liu , James Cheng , Ivor W. Tsang , Lijun Zhang , Dacheng Tao , Licheng Jiao

We study the spectral properties of a class of random matrices of the form $S_n^{-} = n^{-1}(X_1 X_2^* - X_2 X_1^*)$ where $X_k = \Sigma^{1/2}Z_k$, for $k=1,2$, $Z_k$'s are independent $p\times n$ complex-valued random matrices, and…

Statistics Theory · Mathematics 2024-11-27 Javed Hazarika , Debashis Paul

One of the most common methods to train machine learning algorithms today is the stochastic gradient descent (SGD). In a distributed setting, SGD-based algorithms have been shown to converge theoretically under specific circumstances. A…

Machine Learning · Computer Science 2025-08-22 Soumya Sarkar , Shweta Jain

Scaling laws provide compact descriptions of how prediction error varies with compute, model size, and data, but existing theory mainly treats single-sample SGD or full data reuse, leaving the role of mini-batching unclear. We study batch…

Machine Learning · Computer Science 2026-05-26 Ziyan Chen , Ding-Xuan Zhou

Uniform stability is a notion of algorithmic stability that bounds the worst case change in the model output by the algorithm when a single data point in the dataset is replaced. An influential work of Hardt et al. (2016) provides strong…

Machine Learning · Computer Science 2020-06-15 Raef Bassily , Vitaly Feldman , Cristóbal Guzmán , Kunal Talwar

Let A be an n*n random matrix with mean zero and independent inhomogeneous non-constant subgaussian entries. We get that for any k<c\sqrt{n}, the probability of the matrix has a lower rank than n-k that is sub-exponential. Furthermore, we…

Probability · Mathematics 2025-01-28 Guozheng Dai , Zeyan Song , Hanchao Wang

Theoretical works on supervised transfer learning (STL) -- where the learner has access to labeled samples from both source and target distributions -- have for the most part focused on statistical aspects of the problem, while efficient…

Machine Learning · Statistics 2025-07-08 Yuyang Deng , Samory Kpotufe

Let $Q_n$ denote a random symmetric $n$ by $n$ matrix, whose upper diagonal entries are i.i.d. Bernoulli random variables (which take values 0 and 1 with probability 1/2). We prove that $Q_n$ is non-singular with probability…

Probability · Mathematics 2007-05-23 Kevin Costello , Terence Tao , Van Vu

Score-based generative models (SGMs) are powerful tools to sample from complex data distributions. Their underlying idea is to (i) run a forward process for time $T_1$ by adding noise to the data, (ii) estimate its score function, and (iii)…

Machine Learning · Computer Science 2024-06-06 Francesco Pedrotti , Jan Maas , Marco Mondelli