中文
相关论文

相关论文: On the interplay between noise and curvature and i…

200 篇论文

The choice of batch-size in a stochastic optimization algorithm plays a substantial role for both optimization and generalization. Increasing the batch-size used typically improves optimization but degrades generalization. To address the…

机器学习 · 计算机科学 2020-03-03 Yeming Wen , Kevin Luk , Maxime Gazeau , Guodong Zhang , Harris Chan , Jimmy Ba

Accurate state estimation requires careful consideration of uncertainty surrounding the process and measurement models; these characteristics are usually not well-known and need an experienced designer to select the covariance matrices. An…

机器学习 · 统计学 2025-07-18 Pardha Sai Krishna Ala , Ameya Salvi , Venkat Krovi , Matthias Schmid

Standard first-order stochastic optimization algorithms base their updates solely on the average mini-batch gradient, and it has been shown that tracking additional quantities such as the curvature can help de-sensitize common…

机器学习 · 计算机科学 2020-11-11 Ricky T. Q. Chen , Dami Choi , Lukas Balles , David Duvenaud , Philipp Hennig

Second-order information -- such as curvature or data covariance -- is critical for optimisation, diagnostics, and robustness. However, in many modern settings, only the gradients are observable. We show that the gradients alone can reveal…

机器学习 · 计算机科学 2026-04-08 Arash Jamshidi , Katsiaryna Haitsiukevich , Kai Puolamäki

The effect of a small amount of noise on the standard mapping is considered. Whenever the standard mapping possesses accelerator modes (where the action increases approximately linearly with time), the diffusion coefficient contains a term…

混沌动力学 · 物理学 2007-05-23 Charles F. F. Karney , Alexander B. Rechester , Roscoe B. White

Stochastic inverse problems considered in this article consist of estimating the probability distributions of intrinsically random inputs of computer models. These estimations are based on observable outputs affected by model noise, and…

统计理论 · 数学 2025-03-17 Nicolas Bousquet , Mélanie Blazère , Thomas Cerbelaud

We analyze gradient descent with randomly weighted data points in a linear regression model, under a generic weighting distribution. This includes various forms of stochastic gradient descent, importance sampling, but also extends to…

机器学习 · 统计学 2025-12-12 Gabriel Clara , Yazan Mash'al

A relationship between the Fisher information and the characteristic function is established with the help of two inequalities. A necessary and sufficient condition for equality is found. These results are used to determine the asymptotic…

信息论 · 计算机科学 2010-07-12 Cihan Tepedelenlioglu , Mahesh K. Banavar , Andreas Spanias

Decentralized optimization is typically studied under the assumption of noise-free transmission. However, real-world scenarios often involve the presence of noise due to factors such as additive white Gaussian noise channels or…

最优化与控制 · 数学 2023-07-28 Suhail M. Shah , Raghu Bollapragada

We formulate and study a general family of (continuous-time) stochastic dynamics for accelerated first-order minimization of smooth convex functions. Building on an averaging formulation of accelerated mirror descent, we propose a…

最优化与控制 · 数学 2017-07-20 Walid Krichene , Peter L. Bartlett

We introduce novel variants of momentum by incorporating the variance of the stochastic loss function. The variance characterizes the confidence or uncertainty of the local features of the averaged loss surface across the i.i.d. subsets of…

机器学习 · 计算机科学 2019-05-31 Vineeth S. Bhaskara , Sneha Desai

Stochastic gradient algorithms are the main focus of large-scale optimization problems and led to important successes in the recent advancement of the deep learning algorithms. The convergence of SGD depends on the careful choice of…

机器学习 · 计算机科学 2017-03-03 Caglar Gulcehre , Jose Sotelo , Marcin Moczulski , Yoshua Bengio

The objective function of a matrix factorization model usually aims to minimize the average of a regression error contributed by each element. However, given the existence of stochastic noises, the implicit deviations of sample data from…

机器学习 · 计算机科学 2016-10-31 Guang-He Lee , Shao-Wen Yang , Shou-De Lin

Model merging combines independent solutions with different capabilities into a single one while maintaining the same inference cost. Two popular approaches are linear interpolation, which simply averages multiple model weights, and task…

机器学习 · 计算机科学 2026-04-22 Chenxiang Zhang , Alexander Theus , Damien Teney , Antonio Orvieto , Jun Pang , Sjouke Mauw

Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem…

机器学习 · 计算机科学 2019-09-10 Mo Zhou , Tianyi Liu , Yan Li , Dachao Lin , Enlu Zhou , Tuo Zhao

A longstanding debate surrounds the related hypotheses that low-curvature minima generalize better, and that SGD discourages curvature. We offer a more complete and nuanced view in support of both. First, we show that curvature harms test…

机器学习 · 统计学 2022-07-29 Arwen V. Bradley , Carlos Alberto Gomez-Uribe , Manish Reddy Vuyyuru

Under mild assumptions stochastic gradient methods asymptotically achieve an optimal rate of convergence if the arithmetic mean of all iterates is returned as an approximate optimal solution. However, in the absence of stochastic noise, the…

最优化与控制 · 数学 2022-10-06 Melinda Hagedorn , Florian Jarre

Recent studies inspired by results from random matrix theory [1,2,3] found that covariance matrices determined from empirical financial time series appear to contain such a high amount of noise that their structure can essentially be…

统计力学 · 物理学 2009-11-07 Szilard Pafka , Imre Kondor

In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics. Whereas prior work has focused on how different learning…

According to recent findings [1,2], empirical covariance matrices deduced from financial return series contain such a high amount of noise that, apart from a few large eigenvalues and the corresponding eigenvectors, their structure can…

统计力学 · 物理学 2009-11-07 Szilard Pafka , Imre Kondor
‹ 上一页 1 2 3 10 下一页 ›