English
Related papers

Related papers: Can Single-Shuffle SGD be Better than Reshuffling …

200 papers

In this paper, we consider distributed optimization problems where $n$ agents, each possessing a local cost function, collaboratively minimize the average of the local cost functions over a connected network. To solve the problem, we…

Optimization and Control · Mathematics 2023-03-24 Kun Huang , Xiao Li , Andre Milzarek , Shi Pu , Junwen Qiu

We study modeling and inference with the Elliptical Gamma Distribution (EGD). We consider maximum likelihood (ML) estimation for EGD scatter matrices, a task for which we develop new fixed-point algorithms. Our algorithms are efficient and…

Computation · Statistics 2018-06-04 Reshad Hosseini , Suvrit Sra , Lucas Theis , Matthias Bethge

In this paper we improve the best known constant for the discrepancy formulated in the Komlos Conjecture. The result is based on the improvement of the subgaussian bound for the random vector constructed in the Gram-Schmidt Random Walk…

Probability · Mathematics 2024-04-09 Witold Bednorz , Piotr Godlewski

A modification of the generalized shift-splitting (GSS) method is presented for solving singular saddle point problems. In this kind of modification, the diagonal shift matrix is replaced by a block diagonal matrix which is symmetric…

Numerical Analysis · Mathematics 2017-04-26 Davod Khojasteh Salkuyeh , Maryam Rahimian

We investigate matrix models in three dimensions where the global $\text{SU}(N)$ symmetry acts via the adjoint map. Analyzing their ground state which is homogeneous in space and can carry either a unique or multiple fixed charges, we show…

High Energy Physics - Theory · Physics 2018-08-01 Orestis Loukas

The Rayleigh conjecture about convergence up to the boundary of the series representing the scattered field in the exterior of an obstacle $D$ is widely used by engineers in applications. However this conjecture is false for some obstacles.…

Numerical Analysis · Mathematics 2007-05-23 A. G. Ramm , S. Gutman

We explore the asymptotic convergence and nonasymptotic maximal inequalities of supermartingales and backward submartingales in the space of positive semidefinite matrices. These are natural matrix analogs of scalar nonnegative…

Probability · Mathematics 2025-10-21 Hongjian Wang , Aaditya Ramdas

We revisit the classical problem of finding an approximately stationary point of the average of $n$ smooth and possibly nonconvex functions. The optimal complexity of stochastic first-order methods in terms of the number of gradient…

Machine Learning · Computer Science 2022-06-07 Alexander Tyurin , Lukang Sun , Konstantin Burlachenko , Peter Richtárik

Shuffling gradient methods are widely used in modern machine learning tasks and include three popular implementations: Random Reshuffle (RR), Shuffle Once (SO), and Incremental Gradient (IG). Compared to the empirical success, the…

Machine Learning · Computer Science 2024-06-07 Zijian Liu , Zhengyuan Zhou

SGD with momentum (SGDM) has been widely applied in many machine learning tasks, and it is often applied with dynamic stepsizes and momentum weights tuned in a stagewise manner. Despite of its empirical advantage over SGD, the role of…

Optimization and Control · Mathematics 2020-08-19 Yanli Liu , Yuan Gao , Wotao Yin

We investigate a one-time single shelf shuffle by establishing the position matrix explicitly. In some cases, we prove a no-feedback optimal guessing strategy. A general no-feedback strategy is conjectured, and asymptotics for the expected…

Probability · Mathematics 2025-07-15 Alexander Clay

Let $\a$ be a complex random variable with mean zero and bounded variance $\sigma^{2}$. Let $N_{n}$ be a random matrix of order $n$ with entries being i.i.d. copies of $\a$. Let $\lambda_{1}, ..., \lambda_{n}$ be the eigenvalues of…

Probability · Mathematics 2008-02-29 Terence Tao , Van Vu

We consider a random bistochastic matrix of size $n$ of the form $M Q$ where $M$ is a uniformly distributed permutation matrix and $Q$ is a given bistochastic matrix. Under mild sparsity and regularity assumptions on $Q$, we prove that the…

Dynamical Systems · Mathematics 2019-03-26 Charles Bordenave , Yanqi Qiu , Yiwei Zhang

Stochastic gradient descent (SGD) algorithm and its variations have been effectively used to optimize neural network models. However, with the rapid growth of big data and deep learning, SGD is no longer the most suitable choice due to its…

Machine Learning · Computer Science 2024-02-13 Anuraganand Sharma

Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effects of random reshuffling-type methods are fairly well understood in the smooth setting,…

Optimization and Control · Mathematics 2025-07-29 Junwen Qiu , Xiao Li , Andre Milzarek

This article examines the implicit regularization effect of Stochastic Gradient Descent (SGD). We consider the case of SGD without replacement, the variant typically used to optimize large-scale neural networks. We analyze this algorithm in…

Machine Learning · Computer Science 2024-04-23 Pierfrancesco Beneventano

We revisit the use of Stochastic Gradient Descent (SGD) for solving convex optimization problems that serve as highly popular convex relaxations for many important low-rank matrix recovery problems such as \textit{matrix completion},…

Machine Learning · Computer Science 2020-06-16 Dan Garber

Large-scale nonconvex optimization problems are ubiquitous in modern machine learning, and among practitioners interested in solving them, Stochastic Gradient Descent (SGD) reigns supreme. We revisit the analysis of SGD in the nonconvex…

Optimization and Control · Mathematics 2020-07-27 Ahmed Khaled , Peter Richtárik

We provide non-asymptotic, relative deviation bounds for the eigenvalues of empirical covariance and Gram matrices in general settings. Unlike typical uniform bounds, which may fail to capture the behavior of smaller eigenvalues, our…

Probability · Mathematics 2025-05-28 Daniel Barzilai , Ohad Shamir

We examine the use of different randomisation policies for stochastic gradient algorithms used in sampling, based on first-order (or overdamped) Langevin dynamics, the most popular of which is known as Stochastic Gradient Langevin Dynamics.…

Numerical Analysis · Mathematics 2025-12-16 Luke Shaw , Peter A. Whalley
‹ Prev 1 3 4 5 6 7 10 Next ›