Related papers: The Mean-Field Dynamics of Transformers
By the Wolff's cluster Monte Carlo simulations and numerical minimization within a mean field approach, we study the low temperature phase diagram of water, adopting a cell model that reproduces the known properties of water in its fluid…
This paper introduces Generalized Attention Flow (GAF), a novel feature attribution method for Transformer-based models to address the limitations of current approaches. By extending Attention Flow and replacing attention weights with the…
Many complex dynamical systems in the real world, including ecological, climate, financial, and power-grid systems, often show critical transitions, or tipping points, in which the system's dynamics suddenly transit into a qualitatively…
In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime. We identify the set of parameterizations that admit well-defined infinite width and depth limits, allowing the…
Transformer-based deep learning models have achieved state-of-the-art performance across numerous language and vision tasks. While the self-attention mechanism, a core component of transformers, has proven capable of handling complex data…
We study a broad class of high-dimensional mean-field exchange models, encompassing both noisy and singular dynamics, along with their dual processes. This includes a generalized version of the averaging process as well as some…
The phenomenon of Bose-Einstein condensation of dilute gases in traps is reviewed from a theoretical perspective. Mean-field theory provides a framework to understand the main features of the condensation and the role of interactions…
Transformer-based architectures achieve state-of-the-art performance across a wide range of tasks in natural language processing, computer vision, and speech processing. However, their immense capacity often leads to overfitting, especially…
The Hamiltonian Mean Field (HMF) model has a low-energy phase where $N$ particles are trapped inside a cluster. Here, we investigate some properties of the trapping/untrapping mechanism of a single particle into/outside the cluster. Since…
We conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory. We investigate the mechanisms through which different components of Transformer, such as the…
Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine…
We reduce the dynamics of an ensemble of mean-coupled Stuart-Landau oscillators close to the synchronized solution. In particular, we map the system onto the center manifold of the Benjamin-Feir instability, the bifurcation destabilizing…
An analytic framework based on partial differential equations is derived for certain dynamic clustering methods. The proposed mathematical framework is based on the application of the conservation law in physics to characterize successive…
The self-attention mechanism has been a key factor in the advancement of vision Transformers. However, its quadratic complexity imposes a heavy computational burden in high-resolution scenarios, restricting the practical application.…
Linearization of attention using various kernel approximation and kernel learning techniques has shown promise. Past methods used a subset of combinations of component functions and weight matrices within the random feature paradigm. We…
The large amounts of data from molecular biology and neuroscience have lead to a renewed interest in the inverse Ising problem: how to reconstruct parameters of the Ising model (couplings between spins and external fields) from a number of…
Standard attention-based transformers are known to exhibit instability under learning rate overspecification during training, particularly at high learning rates. While various methods have been proposed to improve resilience to such…
Transformers have become a standard neural network architecture for many NLP problems, motivating theoretical analysis of their power in terms of formal languages. Recent work has shown that transformers with hard attention are quite…
We study interacting particle systems of Kuramoto-type. Our focus is on the dynamical relation between the partial differential equation (PDE) arising in the continuum limit (CL) and the one obtained in the mean-field limit (MFL). Both…
The well-posedness of a multi-population dynamical system with an entropy regularization and its convergence to a suitable mean-field approximation are proved, under a general set of assumptions. Under further assumptions on the evolution…