English
Related papers

Related papers: A Mechanism Study of Delayed Loss Spikes in Batch-…

200 papers

This paper is concerned with the problem of regularization by noise of systems of reaction-diffusion equations with mass control. It is known that $\textit{strong}$ solutions to such systems of PDEs may blow-up in finite time. Moreover, for…

Analysis of PDEs · Mathematics 2023-11-30 Antonio Agresti

New methods are developed for the stabilization of a linear system with general time-varying distributed delays existing at the system's states, inputs and outputs. In contrast to most existing literature where the function of time-varying…

Systems and Control · Electrical Eng. & Systems 2024-12-20 Qian Feng , Sing Kiong Nguang , Wilfrid Perruquetti

Due to delays in the adaptation of production or delivery rates, supply chains can be dynamically unstable with respect to perturbations in the consumption rate, which is known as "bull-whip effect". Here, we study several conceivable…

Statistical Mechanics · Physics 2009-11-10 Takashi Nagatani , Dirk Helbing

Normalization techniques such as Batch Normalization have been applied successfully for training deep neural networks. Yet, despite its apparent empirical benefits, the reasons behind the success of Batch Normalization are mostly…

Machine Learning · Statistics 2018-10-09 Jonas Kohler , Hadi Daneshmand , Aurelien Lucchi , Ming Zhou , Klaus Neymeyr , Thomas Hofmann

Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of…

Machine Learning · Statistics 2019-10-29 Ryo Karakida , Shotaro Akaho , Shun-ichi Amari

In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a fundamental question: *does weight decay stabilize training…

Machine Learning · Computer Science 2026-05-19 Marius Saether , Amir Kolic , Tomaso Poggio , Pierfrancesco Beneventano

We study stochastic gradient descent (SGD) for composite optimization problems with $N$ sequential operators subject to perturbations in both the forward and backward passes. Unlike classical analyses that treat gradient noise as additive…

Optimization and Control · Mathematics 2026-02-25 Boao Kong , Hengrui Zhang , Kun Yuan

While many real-world data streams imply that they change frequently in a nonstationary way, most of deep learning methods optimize neural networks on training data, and this leads to severe performance degradation when dataset shift…

Machine Learning · Computer Science 2021-07-02 Wonju Lee , Seok-Yong Byun , Jooeun Kim , Minje Park , Kirill Chechil

We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks. We present empirical observations that commonly used large step sizes (i) lead the iterates to jump from one side of…

Machine Learning · Computer Science 2023-06-08 Maksym Andriushchenko , Aditya Varre , Loucas Pillaud-Vivien , Nicolas Flammarion

We study a special case of the problem of statistical learning without the i.i.d. assumption. Specifically, we suppose a learning method is presented with a sequence of data points, and required to make a prediction (e.g., a classification)…

Machine Learning · Computer Science 2018-05-22 Steve Hanneke , Liu Yang

We propose that the grokking phenomenon, where the train loss of a neural network decreases much earlier than its test loss, can arise due to a neural network transitioning from lazy training dynamics to a rich, feature learning regime. To…

Machine Learning · Statistics 2024-04-12 Tanishq Kumar , Blake Bordelon , Samuel J. Gershman , Cengiz Pehlevan

While synchronized states, and the dynamical pathways through which they emerge, are often regarded as the paradigm to understand the dynamics of information spreading on undirected networks of nonlinear dynamical systems, when we consider…

Adaptation and Self-Organizing Systems · Physics 2025-07-10 Giulio Colombini , Nicola Guglielmi , Armando Bazzani

In real-world networks the interactions between network elements are inherently time-delayed. These time-delays can not only slow the network but can have a destabilizing effect on the network's dynamics leading to poor performance. The…

Optimization and Control · Mathematics 2024-09-23 David Reber , Benjamin Webb

Stochastic Gradient Descent (SGD) based methods have been widely used for training large-scale machine learning models that also generalize well in practice. Several explanations have been offered for this generalization performance, a…

Machine Learning · Computer Science 2021-02-11 Yikai Zhang , Wenjia Zhang , Sammy Bald , Vamsi Pingali , Chao Chen , Mayank Goswami

In this paper we address the issue of output instability of deep neural networks: small perturbations in the visual input can significantly distort the feature embeddings and output of a neural network. Such instability affects many deep…

Computer Vision and Pattern Recognition · Computer Science 2016-04-18 Stephan Zheng , Yang Song , Thomas Leung , Ian Goodfellow

I investigate a stronger form of regularization by deactivating neurons for extended periods, a departure from the temporary changes of methods like Dropout. However, this long-term dynamism introduces a critical challenge: severe training…

Machine Learning · Computer Science 2025-09-26 Zichuan Yang

We conjecture that the inherent difference in generalisation between adaptive and non-adaptive gradient methods in deep learning stems from the increased estimation noise in the flattest directions of the true loss surface. We demonstrate…

Machine Learning · Statistics 2022-03-17 Diego Granziol , Nicholas Baskerville

We consider the modeling, stability analysis and controller design problems for discrete-time LTI systems with state feedback, when the actuation signal is subject to switching propagation delays, due to e.g. the routing in a multi-hop…

Optimization and Control · Mathematics 2014-01-09 R. M. Jungers , A. D'Innocenzo , M. D. Di Benedetto

While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries. We introduce statistically…

Artificial Intelligence · Computer Science 2026-02-17 Yangxinyu Xie , Tao Wang , Soham Mallick , Yan Sun , Georgy Noarov , Mengxin Yu , Tanwi Mallick , Weijie J. Su , Edgar Dobriban

We analyze the training dynamics for deep linear networks using a new metric - layer imbalance - which defines the flatness of a solution. We demonstrate that different regularization methods, such as weight decay or noise data…

Machine Learning · Computer Science 2020-07-21 Boris Ginsburg
‹ Prev 1 4 5 6 7 8 10 Next ›