中文
相关论文

相关论文: Can Single-Shuffle SGD be Better than Reshuffling …

200 篇论文

In this paper, we consider distributed optimization problems where $n$ agents, each possessing a local cost function, collaboratively minimize the average of the local cost functions over a connected network. To solve the problem, we…

最优化与控制 · 数学 2023-03-24 Kun Huang , Xiao Li , Andre Milzarek , Shi Pu , Junwen Qiu

We study modeling and inference with the Elliptical Gamma Distribution (EGD). We consider maximum likelihood (ML) estimation for EGD scatter matrices, a task for which we develop new fixed-point algorithms. Our algorithms are efficient and…

统计计算 · 统计学 2018-06-04 Reshad Hosseini , Suvrit Sra , Lucas Theis , Matthias Bethge

In this paper we improve the best known constant for the discrepancy formulated in the Komlos Conjecture. The result is based on the improvement of the subgaussian bound for the random vector constructed in the Gram-Schmidt Random Walk…

概率论 · 数学 2024-04-09 Witold Bednorz , Piotr Godlewski

A modification of the generalized shift-splitting (GSS) method is presented for solving singular saddle point problems. In this kind of modification, the diagonal shift matrix is replaced by a block diagonal matrix which is symmetric…

数值分析 · 数学 2017-04-26 Davod Khojasteh Salkuyeh , Maryam Rahimian

We investigate matrix models in three dimensions where the global $\text{SU}(N)$ symmetry acts via the adjoint map. Analyzing their ground state which is homogeneous in space and can carry either a unique or multiple fixed charges, we show…

高能物理 - 理论 · 物理学 2018-08-01 Orestis Loukas

The Rayleigh conjecture about convergence up to the boundary of the series representing the scattered field in the exterior of an obstacle $D$ is widely used by engineers in applications. However this conjecture is false for some obstacles.…

数值分析 · 数学 2007-05-23 A. G. Ramm , S. Gutman

We explore the asymptotic convergence and nonasymptotic maximal inequalities of supermartingales and backward submartingales in the space of positive semidefinite matrices. These are natural matrix analogs of scalar nonnegative…

概率论 · 数学 2025-10-21 Hongjian Wang , Aaditya Ramdas

We revisit the classical problem of finding an approximately stationary point of the average of $n$ smooth and possibly nonconvex functions. The optimal complexity of stochastic first-order methods in terms of the number of gradient…

机器学习 · 计算机科学 2022-06-07 Alexander Tyurin , Lukang Sun , Konstantin Burlachenko , Peter Richtárik

Shuffling gradient methods are widely used in modern machine learning tasks and include three popular implementations: Random Reshuffle (RR), Shuffle Once (SO), and Incremental Gradient (IG). Compared to the empirical success, the…

机器学习 · 计算机科学 2024-06-07 Zijian Liu , Zhengyuan Zhou

SGD with momentum (SGDM) has been widely applied in many machine learning tasks, and it is often applied with dynamic stepsizes and momentum weights tuned in a stagewise manner. Despite of its empirical advantage over SGD, the role of…

最优化与控制 · 数学 2020-08-19 Yanli Liu , Yuan Gao , Wotao Yin

We investigate a one-time single shelf shuffle by establishing the position matrix explicitly. In some cases, we prove a no-feedback optimal guessing strategy. A general no-feedback strategy is conjectured, and asymptotics for the expected…

概率论 · 数学 2025-07-15 Alexander Clay

Let $\a$ be a complex random variable with mean zero and bounded variance $\sigma^{2}$. Let $N_{n}$ be a random matrix of order $n$ with entries being i.i.d. copies of $\a$. Let $\lambda_{1}, ..., \lambda_{n}$ be the eigenvalues of…

概率论 · 数学 2008-02-29 Terence Tao , Van Vu

We consider a random bistochastic matrix of size $n$ of the form $M Q$ where $M$ is a uniformly distributed permutation matrix and $Q$ is a given bistochastic matrix. Under mild sparsity and regularity assumptions on $Q$, we prove that the…

动力系统 · 数学 2019-03-26 Charles Bordenave , Yanqi Qiu , Yiwei Zhang

Stochastic gradient descent (SGD) algorithm and its variations have been effectively used to optimize neural network models. However, with the rapid growth of big data and deep learning, SGD is no longer the most suitable choice due to its…

机器学习 · 计算机科学 2024-02-13 Anuraganand Sharma

Random reshuffling techniques are prevalent in large-scale applications, such as training neural networks. While the convergence and acceleration effects of random reshuffling-type methods are fairly well understood in the smooth setting,…

最优化与控制 · 数学 2025-07-29 Junwen Qiu , Xiao Li , Andre Milzarek

This article examines the implicit regularization effect of Stochastic Gradient Descent (SGD). We consider the case of SGD without replacement, the variant typically used to optimize large-scale neural networks. We analyze this algorithm in…

机器学习 · 计算机科学 2024-04-23 Pierfrancesco Beneventano

We revisit the use of Stochastic Gradient Descent (SGD) for solving convex optimization problems that serve as highly popular convex relaxations for many important low-rank matrix recovery problems such as \textit{matrix completion},…

机器学习 · 计算机科学 2020-06-16 Dan Garber

Large-scale nonconvex optimization problems are ubiquitous in modern machine learning, and among practitioners interested in solving them, Stochastic Gradient Descent (SGD) reigns supreme. We revisit the analysis of SGD in the nonconvex…

最优化与控制 · 数学 2020-07-27 Ahmed Khaled , Peter Richtárik

We provide non-asymptotic, relative deviation bounds for the eigenvalues of empirical covariance and Gram matrices in general settings. Unlike typical uniform bounds, which may fail to capture the behavior of smaller eigenvalues, our…

概率论 · 数学 2025-05-28 Daniel Barzilai , Ohad Shamir

We examine the use of different randomisation policies for stochastic gradient algorithms used in sampling, based on first-order (or overdamped) Langevin dynamics, the most popular of which is known as Stochastic Gradient Langevin Dynamics.…

数值分析 · 数学 2025-12-16 Luke Shaw , Peter A. Whalley