English
Related papers

Related papers: Gradients are Not All You Need

200 papers

Regularizing the gradient norm of the output of a neural network with respect to its inputs is a powerful technique, rediscovered several times. This paper presents evidence that gradient regularization can consistently improve…

Machine Learning · Computer Science 2018-05-28 Dániel Varga , Adrián Csiszárik , Zsolt Zombori

Although much progress has been made towards robust deep learning, a significant gap in robustness remains between real-world perturbations and more narrowly defined sets typically studied in adversarial defenses. In this paper, we aim to…

Machine Learning · Computer Science 2020-10-09 Eric Wong , J. Zico Kolter

Recurrent Neural Networks (RNNs) frequently exhibit complicated dynamics, and their sensitivity to the initialization process often renders them notoriously hard to train. Recent works have shed light on such phenomena analyzing when…

Machine Learning · Computer Science 2022-10-12 Vaggos Chatziafratis , Ioannis Panageas , Clayton Sanford , Stelios Andrew Stavroulakis

After a more than decade-long period of relatively little research activity in the area of recurrent neural networks, several new developments will be reviewed here that have allowed substantial progress both in understanding and in…

Machine Learning · Computer Science 2012-12-17 Yoshua Bengio , Nicolas Boulanger-Lewandowski , Razvan Pascanu

We conjecture that the inherent difference in generalisation between adaptive and non-adaptive gradient methods in deep learning stems from the increased estimation noise in the flattest directions of the true loss surface. We demonstrate…

Machine Learning · Statistics 2022-03-17 Diego Granziol , Nicholas Baskerville

Partitioning a set of elements into an unknown number of mutually exclusive subsets is essential in many machine learning problems. However, assigning elements, such as samples in a dataset or neurons in a network layer, to an unknown and…

Machine Learning · Computer Science 2023-11-10 Thomas M. Sutter , Alain Ryser , Joram Liebeskind , Julia E. Vogt

Many real-world combinatorial problems involve uncertain parameters, which can be predicted given contextual features and historical data. These `predict-then-optimize' or `contextual optimization' problems have gained significant…

Machine Learning · Computer Science 2026-05-19 Noah Schutte , Senne Berden , Tias Guns , Krzysztof Postek , Neil Yorke-Smith

We undertake a systematic study of the dynamics of Boolean networks to determine the origin of chaos observed in recent experiments. Networks with nodes consisting of ideal logic gates are known to display either steady states, periodic…

Chaotic Dynamics · Physics 2015-05-14 Hugo L. D. de S. Cavalcante , Daniel J. Gauthier , Joshua E. S. Socolar , Rui Zhang

We give a rigorous analysis of the statistical behavior of gradients in a randomly initialized fully connected network N with ReLU activations. Our results show that the empirical variance of the squares of the entries in the input-output…

Machine Learning · Statistics 2018-10-30 Boris Hanin

The Hopfield network has been applied to solve optimization problems over decades. However, it still has many limitations in accomplishing this task. Most of them are inherited from the optimization algorithms it implements. The computation…

Neural and Evolutionary Computing · Computer Science 2007-05-23 Xiaofei Huang

Algorithm unrolling is ubiquitous in machine learning, particularly in hyperparameter optimization and meta-learning, where Jacobians of solution mappings are computed by differentiating through iterative algorithms. Although unrolling is…

Machine Learning · Computer Science 2026-02-24 Sheheryar Mehmood , Florian Knoll , Peter Ochs

We characterize a quantum neural network's error in terms of the network's scrambling properties via the out-of-time-ordered correlator. A network can be trained by optimizing either a loss function or a cost function. We show that, with…

Quantum Physics · Physics 2022-03-23 Roy J. Garcia , Kaifeng Bu , Arthur Jaffe

We observe a novel 'multiple-descent' phenomenon during the training process of LSTM, in which the test loss goes through long cycles of up and down trend multiple times after the model is overtrained. By carrying out asymptotic stability…

Machine Learning · Computer Science 2025-05-27 Wenbo Wei , Nicholas Chong Jia Le , Choy Heng Lai , Ling Feng

Differentiable programming has recently received much interest as a paradigm that facilitates taking gradients of computer programs. While the corresponding flexible gradient-based optimization approaches so far have been used predominantly…

The deep learning revolution has spurred a rise in advances of using AI in sciences. Within physical sciences the main focus has been on discovery of dynamical systems from observational data. Yet the reliability of learned surrogates and…

Dynamical Systems · Mathematics 2025-11-13 Zakhar Shumaylov , Peter Zaika , Philipp Scholl , Gitta Kutyniok , Lior Horesh , Carola-Bibiane Schönlieb

This paper addresses the efficient computation of Jacobian matrices for programs composed of sequential differentiable subprograms. By representing the overall Jacobian as a chain product of the Jacobians of these subprograms, we reduce the…

Discrete Mathematics · Computer Science 2025-05-12 Simon Märtens , Uwe Naumann

Continual learning is an emerging paradigm in machine learning, wherein a model is exposed in an online fashion to data from multiple different distributions (i.e. environments), and is expected to adapt to the distribution change.…

Machine Learning · Computer Science 2022-03-29 Binghui Peng , Andrej Risteski

Chaotic oscillators have gained significant attention in the research community because of their ability to reproduce and investigate the complex dynamics of real-world phenomena. Recent advances in the design of chaotic oscillator…

Chaotic Dynamics · Physics 2026-03-19 Toni Ivas , Georgios Violakis , Roland Richter , Patrik Hoffmann , Sergey Shevchik

Training deep neural networks is a highly nontrivial task, involving carefully selecting appropriate training algorithms, scheduling step sizes and tuning other hyperparameters. Trying different combinations can be quite labor-intensive and…

Machine Learning · Computer Science 2017-06-13 Kaifeng Lv , Shunhua Jiang , Jian Li

The relative power of quantum algorithms, using an adaptive access to quantum devices, versus classical post-processing methods that rely only on an initial quantum data set, remains the subject of active debate. Here, we present evidence…

Quantum Physics · Physics 2025-10-02 Oleksandr Kyriienko , Chukwudubem Umeano , Zoë Holmes
‹ Prev 1 3 4 5 6 7 10 Next ›