中文
相关论文

相关论文: Can Single-Shuffle SGD be Better than Reshuffling …

200 篇论文

Given a sequence $(M_{n},Q_{n})_{n\ge 1}$ of i.i.d. random variables with generic copy $(M,Q)$ such that $M$ is a regular $d\times d$ matrix and $Q$ takes values in $\mathbb{R}^{d}$, we consider the random difference equation (RDE)…

概率论 · 数学 2013-04-08 Gerold Alsmeyer , Sebastian Mentemeier

In this paper we identify a significant deficiency in the literature on the application of the Relative Gain Array (RGA) formalism in the case of singular matrices. Specifically, we show that the conventional use of the Moore-Penrose…

系统与控制 · 计算机科学 2019-03-06 Jeffrey Uhlmann

An important open problem is the theoretically feasible acceleration of mini-batch SGD-type algorithms on quadratic problems with power-law spectrum. In the non-stochastic setting, the optimal exponent $\xi$ in the loss convergence $L_t\sim…

机器学习 · 计算机科学 2025-03-11 Dmitry Yarotsky , Maksim Velikanov

Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing…

概率论 · 数学 2026-05-11 Bartosz Glowacki , Rafal Kulik , Philippe Soulier

In this work, we consider minimizing the average of a very large number of smooth and possibly non-convex functions, and we focus on two widely used minibatch frameworks to tackle this optimization problem: Incremental Gradient (IG) and…

最优化与控制 · 数学 2024-05-22 Ruggiero Seccia , Corrado Coppola , Giampaolo Liuzzi , Laura Palagi

In this paper we introduce the algorithm and the fixed point hardware to calculate the normalized singular value decomposition of a non-symmetric matrices using Givens fast (approximate) rotations. This algorithm only uses the basic…

数值分析 · 计算机科学 2017-07-18 Ehsan Rohani , Gwan Choi , Mi Lu

We prove that the (non-symmetric) adjacency matrix of a uniform random $d$-regular directed graph on $n$ vertices is asymptotically almost surely invertible, assuming $\min(d,n-d)\ge C\log^2n$ for a sufficiently large constant $C>0$. The…

概率论 · 数学 2015-11-10 Nicholas A. Cook

This paper revisits the convergence of Stochastic Mirror Descent (SMD) in the contemporary nonconvex optimization setting. Existing results for batch-free nonconvex SMD restrict the choice of the distance generating function (DGF) to be…

最优化与控制 · 数学 2024-02-28 Ilyas Fatkhullin , Niao He

We combine two advanced ideas widely used in optimization for machine learning: shuffling strategy and momentum technique to develop a novel shuffling gradient-based method with momentum, coined Shuffling Momentum Gradient (SMG), for…

最优化与控制 · 数学 2021-06-10 Trang H. Tran , Lam M. Nguyen , Quoc Tran-Dinh

We reconsider randomized algorithms for the low-rank approximation of symmetric positive semi-definite (SPSD) matrices such as Laplacian and kernel matrices that arise in data analysis and machine learning applications. Our main results…

机器学习 · 计算机科学 2013-06-05 Alex Gittens , Michael W. Mahoney

The matrix completion problem seeks to recover a $d\times d$ ground truth matrix of low rank $r\ll d$ from observations of its individual elements. Real-world matrix completion is often a huge-scale optimization problem, with $d$ so large…

机器学习 · 计算机科学 2022-10-25 Gavin Zhang , Hong-Ming Chiu , Richard Y. Zhang

A family of symmetric matrices $A_1,\ldots, A_d$ is SDC (simultaneous diagonalization by congruence, also called non-orthogonal joint diagonalization) if there is an invertible matrix $X$ such that every $X^T A_k X$ is diagonal. In this…

数值分析 · 数学 2025-04-30 Haoze He , Daniel Kressner

It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve…

信息论 · 计算机科学 2023-07-24 Rishabh Dudeja , Subhabrata Sen , Yue M. Lu

Decentralized stochastic optimization methods have gained a lot of attention recently, mainly because of their cheap per iteration cost, data locality, and their communication-efficiency. In this paper we introduce a unified convergence…

机器学习 · 计算机科学 2021-03-03 Anastasia Koloskova , Nicolas Loizou , Sadra Boreiri , Martin Jaggi , Sebastian U. Stich

We address the component-based regularisation of a multivariate Generalized Linear Mixed Model (GLMM). A set of random responses Y is modelled by a GLMM, using a set X of explanatory variables, a set T of additional covariates, and random…

统计方法学 · 统计学 2019-08-22 Jocelyn Chauvet , Catherine Trottier , Xavier Bry , Frederic Mortier

Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD. Despite this success, theoretically proving the dominance of local SGD in settings with…

While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we…

机器学习 · 计算机科学 2024-01-17 Ali Ramezani-Kebrya , Kimon Antonakopoulos , Volkan Cevher , Ashish Khisti , Ben Liang

In this paper we propose a new scaling method to study the Schur complements of $SDD_{1}$ matrices. Its core is related to the non-negative property of the inverse $M$-matrix, while numerically improving the Quotient formula. Based on the…

数值分析 · 数学 2025-04-22 Yang Hu , Jianzhou Liu , Wenlong Zeng

We study Deligne's conjecture on the monodromy weight filtration on the nearby cycles in the mixed characteristic case, and reduce it to the nondegeneracy of certain pairings in the semistable case. We also prove a related conjecture of…

代数几何 · 数学 2007-05-23 Morihiko Saito

In this paper we introduce a unified analysis of a large family of variants of proximal stochastic gradient descent ({\tt SGD}) which so far have required different intuitions, convergence analyses, have different applications, and which…

最优化与控制 · 数学 2019-05-28 Eduard Gorbunov , Filip Hanzely , Peter Richtárik