English
Related papers

Related papers: Can Single-Shuffle SGD be Better than Reshuffling …

200 papers

Given a sequence $(M_{n},Q_{n})_{n\ge 1}$ of i.i.d. random variables with generic copy $(M,Q)$ such that $M$ is a regular $d\times d$ matrix and $Q$ takes values in $\mathbb{R}^{d}$, we consider the random difference equation (RDE)…

Probability · Mathematics 2013-04-08 Gerold Alsmeyer , Sebastian Mentemeier

In this paper we identify a significant deficiency in the literature on the application of the Relative Gain Array (RGA) formalism in the case of singular matrices. Specifically, we show that the conventional use of the Moore-Penrose…

Systems and Control · Computer Science 2019-03-06 Jeffrey Uhlmann

An important open problem is the theoretically feasible acceleration of mini-batch SGD-type algorithms on quadratic problems with power-law spectrum. In the non-stochastic setting, the optimal exponent $\xi$ in the loss convergence $L_t\sim…

Machine Learning · Computer Science 2025-03-11 Dmitry Yarotsky , Maksim Velikanov

Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing…

Probability · Mathematics 2026-05-11 Bartosz Glowacki , Rafal Kulik , Philippe Soulier

In this work, we consider minimizing the average of a very large number of smooth and possibly non-convex functions, and we focus on two widely used minibatch frameworks to tackle this optimization problem: Incremental Gradient (IG) and…

Optimization and Control · Mathematics 2024-05-22 Ruggiero Seccia , Corrado Coppola , Giampaolo Liuzzi , Laura Palagi

In this paper we introduce the algorithm and the fixed point hardware to calculate the normalized singular value decomposition of a non-symmetric matrices using Givens fast (approximate) rotations. This algorithm only uses the basic…

Numerical Analysis · Computer Science 2017-07-18 Ehsan Rohani , Gwan Choi , Mi Lu

We prove that the (non-symmetric) adjacency matrix of a uniform random $d$-regular directed graph on $n$ vertices is asymptotically almost surely invertible, assuming $\min(d,n-d)\ge C\log^2n$ for a sufficiently large constant $C>0$. The…

Probability · Mathematics 2015-11-10 Nicholas A. Cook

This paper revisits the convergence of Stochastic Mirror Descent (SMD) in the contemporary nonconvex optimization setting. Existing results for batch-free nonconvex SMD restrict the choice of the distance generating function (DGF) to be…

Optimization and Control · Mathematics 2024-02-28 Ilyas Fatkhullin , Niao He

We combine two advanced ideas widely used in optimization for machine learning: shuffling strategy and momentum technique to develop a novel shuffling gradient-based method with momentum, coined Shuffling Momentum Gradient (SMG), for…

Optimization and Control · Mathematics 2021-06-10 Trang H. Tran , Lam M. Nguyen , Quoc Tran-Dinh

We reconsider randomized algorithms for the low-rank approximation of symmetric positive semi-definite (SPSD) matrices such as Laplacian and kernel matrices that arise in data analysis and machine learning applications. Our main results…

Machine Learning · Computer Science 2013-06-05 Alex Gittens , Michael W. Mahoney

The matrix completion problem seeks to recover a $d\times d$ ground truth matrix of low rank $r\ll d$ from observations of its individual elements. Real-world matrix completion is often a huge-scale optimization problem, with $d$ so large…

Machine Learning · Computer Science 2022-10-25 Gavin Zhang , Hong-Ming Chiu , Richard Y. Zhang

A family of symmetric matrices $A_1,\ldots, A_d$ is SDC (simultaneous diagonalization by congruence, also called non-orthogonal joint diagonalization) if there is an invertible matrix $X$ such that every $X^T A_k X$ is diagonal. In this…

Numerical Analysis · Mathematics 2025-04-30 Haoze He , Daniel Kressner

It has been observed that the performances of many high-dimensional estimation problems are universal with respect to underlying sensing (or design) matrices. Specifically, matrices with markedly different constructions seem to achieve…

Information Theory · Computer Science 2023-07-24 Rishabh Dudeja , Subhabrata Sen , Yue M. Lu

Decentralized stochastic optimization methods have gained a lot of attention recently, mainly because of their cheap per iteration cost, data locality, and their communication-efficiency. In this paper we introduce a unified convergence…

Machine Learning · Computer Science 2021-03-03 Anastasia Koloskova , Nicolas Loizou , Sadra Boreiri , Martin Jaggi , Sebastian U. Stich

We address the component-based regularisation of a multivariate Generalized Linear Mixed Model (GLMM). A set of random responses Y is modelled by a GLMM, using a set X of explanatory variables, a set T of additional covariates, and random…

Methodology · Statistics 2019-08-22 Jocelyn Chauvet , Catherine Trottier , Xavier Bry , Frederic Mortier

Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD. Despite this success, theoretically proving the dominance of local SGD in settings with…

While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we…

Machine Learning · Computer Science 2024-01-17 Ali Ramezani-Kebrya , Kimon Antonakopoulos , Volkan Cevher , Ashish Khisti , Ben Liang

In this paper we propose a new scaling method to study the Schur complements of $SDD_{1}$ matrices. Its core is related to the non-negative property of the inverse $M$-matrix, while numerically improving the Quotient formula. Based on the…

Numerical Analysis · Mathematics 2025-04-22 Yang Hu , Jianzhou Liu , Wenlong Zeng

We study Deligne's conjecture on the monodromy weight filtration on the nearby cycles in the mixed characteristic case, and reduce it to the nondegeneracy of certain pairings in the semistable case. We also prove a related conjecture of…

Algebraic Geometry · Mathematics 2007-05-23 Morihiko Saito

In this paper we introduce a unified analysis of a large family of variants of proximal stochastic gradient descent ({\tt SGD}) which so far have required different intuitions, convergence analyses, have different applications, and which…

Optimization and Control · Mathematics 2019-05-28 Eduard Gorbunov , Filip Hanzely , Peter Richtárik