English
Related papers

Related papers: Subcritical Signal Propagation at Initialization i…

200 papers

Graph Neural Networks (GNNs) often suffer from performance degradation as the network depth increases. This paper addresses this issue by introducing initialization methods that enhance signal propagation (SP) within GNNs. We propose three…

Machine Learning · Computer Science 2025-07-16 Senmiao Wang , Yupeng Chen , Yushun Zhang , Ruoyu Sun , Tian Ding

This paper investigates the key role of Feed-Forward Networks (FFNs) in transformer models by utilizing the Parallel Attention and Feed-Forward Net Design (PAF) architecture, and comparing it to their Series Attention and Feed-Forward Net…

Computation and Language · Computer Science 2023-05-26 Shashank Sonkar , Richard G. Baraniuk

This paper investigates distributed resource allocation optimization over directed graphs with limited communication bandwidth. We develop a novel distributed algorithm that integrates the centralized Proximal Jacobian Alternating Direction…

Optimization and Control · Mathematics 2026-04-17 Xu Du , Boyu Han , Ivano Notarnicola , Karl H. Johansson , Apostolos I. Rikos

Batch Normalization is a key component in almost all state-of-the-art image classifiers, but it also introduces practical challenges: it breaks the independence between training examples within a batch, can incur compute and memory…

Machine Learning · Computer Science 2021-01-28 Andrew Brock , Soham De , Samuel L. Smith

We investigate the learning dynamics of fully-connected neural networks through the lens of gradient signal-to-noise ratio (SNR), examining the behavior of first-order optimizers like Adam in non-convex objectives. By interpreting the…

Machine Learning · Computer Science 2024-03-28 Sokratis J. Anagnostopoulos , Juan Diego Toscano , Nikolaos Stergiopulos , George Em Karniadakis

Adversarial examples are crafted with imperceptible perturbations with the intent to fool neural networks. Against such attacks, adversarial training and its variants stand as the strongest defense to date. Previous studies have pointed out…

Computer Vision and Pattern Recognition · Computer Science 2020-01-30 Alvin Chan , Yi Tay , Yew Soon Ong , Jie Fu

Transformers have become foundational architectures for both natural language and computer vision tasks. However, the high computational cost makes it quite challenging to deploy on resource-constraint devices. This paper investigates the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jialong Guo , Xinghao Chen , Yehui Tang , Yunhe Wang

Pre-propagation graph neural networks (PPGNNs) decouple node feature propagation from transformation: graph diffusion is performed once as preprocessing, and training reduces to dense per-node transformations. This design enables mini-batch…

Machine Learning · Computer Science 2026-05-26 Zichao Yue , Zhiru Zhang

We address the problem of a front propagation in chains with a bi-stable nondegenerate on-site potential and a nonlinear gradient coupling. For a generic nonlinear coupling, one encounters a special regime of transitions, characterized by…

Pattern Formation and Solitons · Physics 2018-03-14 I. B. Shiroky , O. V. Gendelman

Recently popularized graph neural networks achieve the state-of-the-art accuracy on a number of standard benchmark datasets for graph-based semi-supervised learning, improving significantly over existing approaches. These architectures…

Machine Learning · Statistics 2018-03-13 Kiran K. Thekumparampil , Chong Wang , Sewoong Oh , Li-Jia Li

Network pruning is a promising avenue for compressing deep neural networks. A typical approach to pruning starts by training a model and then removing redundant parameters while minimizing the impact on what is learned. Alternatively, a…

Machine Learning · Computer Science 2020-02-18 Namhoon Lee , Thalaiyasingam Ajanthan , Stephen Gould , Philip H. S. Torr

Length Generalization is the essential capacity of autonomous agents to perform tasks in longer contexts than those encountered during training. To systematically study this feat, we test how well models can approximate the next token…

Many deep Convolutional Neural Networks (CNN) make incorrect predictions on adversarial samples obtained by imperceptible perturbations of clean samples. We hypothesize that this is caused by a failure to suppress unusual signals within…

Computer Vision and Pattern Recognition · Computer Science 2016-03-18 Qiyang Zhao , Lewis D Griffin

We investigate forward signal propagation and gradient back propagation in deep, randomly initialized transformers, yielding simple necessary and sufficient conditions on initialization hyperparameters that ensure trainability of deep…

Disordered Systems and Neural Networks · Physics 2024-03-06 Aditya Cowsik , Tamra Nebabu , Xiao-Liang Qi , Surya Ganguli

Dissipation is a ubiquitous phenomenon in dynamical systems encountered in nature because no finite system is fully isolated from its environment. In optical systems, a key challenge facing any technological application has traditionally…

Optics · Physics 2014-10-20 Konstantinos G. Makris , Li Ge , Hakan E. Tureci

Gradient descent during the learning process of a neural network can be subject to many instabilities. The spectral density of the Jacobian is a key component for analyzing stability. Following the works of Pennington et al., such Jacobians…

Machine Learning · Statistics 2023-04-26 Reda Chhaibi , Tariq Daouda , Ezechiel Kahn

Deep networks are vulnerable to adversarial examples. Adversarial Training (AT) has been a standard foundation of modern adversarial defense approaches due to its remarkable effectiveness. However, AT is extremely time-consuming, refraining…

Machine Learning · Computer Science 2024-05-28 Shao-Yuan Lo , Vishal M. Patel

There has been an increasing interest in learning dynamics simulators for model-based control. Compared with off-the-shelf physics engines, a learnable simulator can quickly adapt to unseen objects, scenes, and tasks. However, existing…

Artificial Intelligence · Computer Science 2019-04-19 Yunzhu Li , Jiajun Wu , Jun-Yan Zhu , Joshua B. Tenenbaum , Antonio Torralba , Russ Tedrake

A well-conditioned Jacobian spectrum has a vital role in preventing exploding or vanishing gradients and speeding up learning of deep neural networks. Free probability theory helps us to understand and handle the Jacobian spectrum. We…

Probability · Mathematics 2020-02-13 Tomohiro Hayase

We bring together three key amplification mechanisms in linear dynamical systems: spectral criticality, resonance, and non-normality. We present a unified linear framework that both distinguishes and quantitatively links these effects…

Chaotic Dynamics · Physics 2025-08-18 Virgile Troude , Didier Sornette