中文
相关论文

相关论文: Can Single-Shuffle SGD be Better than Reshuffling …

200 篇论文

When iteratively solving linear systems By=b with Hermitian positive semi-definite $B$, and in particular when solving least-squares problems for $Ax=b$ by reformulating them as $AA^\ast y=b$, it is often observed that SOR-type methods…

数值分析 · 数学 2016-07-21 Peter Oswald , Weiqi Zhou

We study nonconvex finite-sum problems and analyze stochastic variance reduced gradient (SVRG) methods for them. SVRG and related methods have recently surged into prominence for convex optimization given their edge over stochastic gradient…

最优化与控制 · 数学 2016-04-06 Sashank J. Reddi , Ahmed Hefny , Suvrit Sra , Barnabas Poczos , Alex Smola

Stochastic gradient descent (SGD) is a workhorse algorithm for solving large-scale optimization problems in data science and machine learning. Understanding the convergence of SGD is hence of fundamental importance. In this work we examine…

数值分析 · 数学 2024-12-11 Lehan Chen , Yuji Nakatsukasa

Consider an n by n array of cards shuffled in the following manner. An element x of the array is chosen uniformly at random; Then with probability 1/2 the rectangle of cards above and to the left of x is rotated 180 degrees, and with…

概率论 · 数学 2007-05-23 Robin Pemantle

Over the past few years, trace regression models have received considerable attention in the context of matrix completion, quantum state tomography, and compressed sensing. Estimation of the underlying matrix from regularization-based…

机器学习 · 统计学 2015-04-24 Martin Slawski , Ping Li , Matthias Hein

Shuffling-type gradient methods are favored in practice for their simplicity and rapid empirical performance. Despite extensive development of convergence guarantees under various assumptions in recent years, most require the Lipschitz…

机器学习 · 计算机科学 2025-07-15 Qi He , Peiran Yu , Ziyi Chen , Heng Huang

We investigate the emergence of single spin asymmetries (SSA) in hard processes using transverse momentum dependent (TMD) distribution and fragmentation functions. Specifically, the description of SSA involves time reversal-odd functions.…

高能物理 - 唯象学 · 物理学 2008-06-30 P. J. Mulders

1. A standard Gaussian random matrix has full rank with probability 1 and is well-conditioned with a probability quite close to 1 and converging to 1 fast as the matrix deviates from square shape and becomes more rectangular. 2. If we…

数值分析 · 数学 2016-03-17 Victor Y. Pan , Liang Zhao

Despite more than 40 years of research in condensed-matter physics, state-of-the-art approaches for simulating the radial distribution function (RDF) g(r) still rely on binning pair-separations into a histogram. Such methods suffer from…

材料科学 · 物理学 2016-09-05 Thomas W. Rosch , Paul N. Patrone

In this paper, we extend the rectangular side of the shuffle conjecture by stating a rectangular analogue of the square paths conjecture. In addition, we describe a set of combinatorial objects and one statistic that are a first step…

组合数学 · 数学 2023-12-07 Alessandro Iraci , Roberto Pagaria , Giovanni Paolini , Anna Vanden Wyngaerd

The theory part of this paper is sketched as follows. Based on column stochastic average matrix $T_n$ selected as a basic substitution matrix, the method of advanced successive difference substitution is established. Then, a set of…

符号计算 · 计算机科学 2010-04-05 Yong Yao

Understanding the limitations of gradient methods, and stochastic gradient descent (SGD) in particular, is a central challenge in learning theory. To that end, a commonly used tool is the Statistical Queries (SQ) framework, which studies…

机器学习 · 计算机科学 2026-02-06 Daniel Barzilai , Ohad Shamir

In this work, we provide a fundamental unified convergence theorem used for deriving expected and almost sure convergence results for a series of stochastic optimization methods. Our unified theorem only requires to verify several…

最优化与控制 · 数学 2022-10-20 Xiao Li , Andre Milzarek

LocalSGD and SCAFFOLD are widely used methods in distributed stochastic optimization, with numerous applications in machine learning, large-scale data processing, and federated learning. However, rigorously establishing their theoretical…

最优化与控制 · 数学 2025-02-25 Ruichen Luo , Sebastian U Stich , Samuel Horváth , Martin Takáč

The proximal stochastic gradient method (PSGD) is one of the state-of-the-art approaches for stochastic composite-type problems. In contrast to its deterministic counterpart, PSGD has been found to have difficulties with the correct…

最优化与控制 · 数学 2026-03-04 Junwen Qiu , Li Jiang , Andre Milzarek

Patterned random matrices such as the reverse circulant, the symmetric circulant, the Toeplitz and the Hankel matrices and their almost sure limiting spectral distribution (LSD), have attracted much attention. Under the assumption that the…

概率论 · 数学 2022-03-14 Arup Bose , Koushik Saha , Priyanka Sen

Most prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize uniform…

机器学习 · 计算机科学 2023-06-07 Rudrajit Das , Satyen Kale , Zheng Xu , Tong Zhang , Sujay Sanghavi

A few matrix-vector multiplications with random vectors are often sufficient to obtain reasonably good estimates for the norm of a general matrix or the trace of a symmetric positive semi-definite matrix. Several such probabilistic…

数值分析 · 数学 2020-08-11 Zvonimir Bujanović , Daniel Kressner

The switch Markov chain has been extensively studied as the most natural Markov Chain Monte Carlo approach for sampling graphs with prescribed degree sequences. We use comparison arguments with other, less natural but simpler to analyze,…

离散数学 · 计算机科学 2018-10-29 Georgios Amanatidis , Pieter Kleer

The direct sampling method (DSM) has been introduced for non-iterative imaging of small inhomogeneities and is known to be fast, robust, and effective for inverse scattering problems. However, to the best of our knowledge, a full analysis…

数值分析 · 数学 2018-09-26 Sangwoo Kang , Marc Lambert , Won-Kwang Park
‹ 上一页 1 8 9 10 下一页 ›