English
Related papers

Related papers: Recurrent Self-Attention Dynamics: An Energy-Agnos…

200 papers

This work introduces the concept of tangent space regularization for neural-network models of dynamical systems. The tangent space to the dynamics function of many physical systems of interest in control applications exhibits useful…

Machine Learning · Computer Science 2018-06-27 Fredrik Bagge Carlson , Rolf Johansson , Anders Robertsson

Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation slows and diversifies inhibitory timescales, leading to improved task performance that is…

Neurons and Cognition · Quantitative Biology 2025-12-19 Thiparat Chotibut , Oleg Evnin , Weerawit Horinouchi

Analytical treatments of far-from-equilibrium quantum dynamics are few, even in well-thermalizing systems. The celebrated eigenstate thermalization hypothesis (ETH) provides a post hoc ansatz for the matrix elements of observables in the…

Quantum Physics · Physics 2025-10-06 Dominik Hahn , David M. Long , Marin Bukov , Anushya Chandran

Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, recent work has highlighted the inherent instability of SSA…

Machine Learning · Computer Science 2026-02-06 Leo Zhang , James Martens

We reformulate the zero-dimensional hermitean one-matrix model as a (nonlocal) collective field theory, for finite~$N$. The Jacobian arising by changing variables from matrix eigenvalues to their density distribution is treated {\it…

High Energy Physics - Theory · Physics 2010-11-01 Olaf Lechtenfeld

The self-attention (SA) mechanism has demonstrated superior performance across various domains, yet it suffers from substantial complexity during both training and inference. The next-generation architecture, aiming at retaining the…

Machine Learning · Computer Science 2025-01-13 Guoxin Feng

We present a theoretical analysis of the Jacobian of an attention block within a transformer, showing that it is governed by the query, key, and value projections that define the attention mechanism. Leveraging this insight, we introduce a…

Machine Learning · Computer Science 2026-03-10 Hemanth Saratchandran , Simon Lucey

A commonly used approach to study stability in a complex system is by analyzing the Jacobian matrix at an equilibrium point of a dynamical system. The equilibrium point is stable if all eigenvalues have negative real parts. Here, by…

Populations and Evolution · Quantitative Biology 2016-09-02 James P. L. Tan

Quantifying interaction strengths between state variables in dynamical systems is essential for understanding ecological networks. Within the empirical dynamic modeling approach, multivariate S-map infers the interaction Jacobian from time…

Populations and Evolution · Quantitative Biology 2024-11-15 Takeshi Miki , Chun-Wei Chang , Po-Ju Ke , Arndt Telschow , Cheng-Han Tsai , Masayuki Ushio , Chih-hao Hsieh

Relaxation dynamics in reversible catalytic reaction networks is studied, revealing two salient behaviors that are reminiscent of glassy behavior: slow relaxation with log(time) dependence of the correlation function, and emergence of a few…

Disordered Systems and Neural Networks · Physics 2015-05-13 Akinori Awazu , Kunihiko Kaneko

Self-attention is essential to Transformer architectures, yet how information is embedded in the self-attention matrices and how different objective functions impact this process remains unclear. We present a mathematical framework to…

Machine Learning · Computer Science 2025-06-04 Matteo Saponati , Pascal Sager , Pau Vilimelis Aceituno , Thilo Stadelmann , Benjamin Grewe

We study signal propagation at initialization in transformers through the averaged partial Jacobian norm (APJN), a measure of gradient amplification across layers. We extend APJN analysis to transformers with bidirectional attention and…

Machine Learning · Computer Science 2026-05-08 Sergey Alekseev

This article provides a comprehensive understanding of optimization in deep learning, with a primary focus on the challenges of gradient vanishing and gradient exploding, which normally lead to diminished model representational ability and…

Machine Learning · Computer Science 2023-11-14 Xianbiao Qi , Jianan Wang , Lei Zhang

We present a mathematical analysis of the effects of Hebbian learning in random recurrent neural networks, with a generic Hebbian learning rule including passive forgetting and different time scales for neuronal activity and learning…

Chaotic Dynamics · Physics 2008-04-07 Benoit Siri , Hugues Berry , Bruno Cessac , Bruno Delord , Mathias Quoy

Transformers are one of the most successful architectures of modern neural networks. At their core there is the so-called attention mechanism, which recently interested the physics community as it can be written as the derivative of an…

Machine Learning · Computer Science 2024-09-25 Francesco D'Amico , Matteo Negri

Self-attention (SA) has become the cornerstone of modern vision backbones for its powerful expressivity over traditional Convolutions (Conv). However, its quadratic complexity remains a critical bottleneck for practical applications. Given…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Hao Yu , Haoyu Chen , Yan Jiang , Wei Peng , Zhaodong Sun , Samuel Kaski , Guoying Zhao

We introduce Robust Filter Attention (RFA), a formulation of self-attention as a robust state estimator. Each token is treated as a noisy observation of a latent trajectory governed by a linear stochastic differential equation (SDE), and…

Machine Learning · Computer Science 2026-05-26 Peter Racioppo

The self-attention mechanism prevails in modern machine learning. It has an interesting functionality of adaptively selecting tokens from an input sequence by modulating the degree of attention localization, which many researchers speculate…

Machine Learning · Statistics 2024-02-06 Han Bao , Ryuichiro Hataya , Ryo Karakida

To make predictions or design control, information on local sensitivity of initial conditions and state-space contraction is both central, and often instrumental. However, it is not always simple to reliably determine instability fields or…

Despite powering modern AI, transformers remain mysteriously brittle to train. We develop a stability theory that explains why pre-LayerNorm works, why DeepNorm uses $N^{-1/4}$ scaling, and why warmup is necessary, all from first…

Machine Learning · Computer Science 2026-02-24 Seyed Morteza Emadi
‹ Prev 1 2 3 10 Next ›