Related papers: Dynamic metastability in the self-attention model
The great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information from the input. Prior work strives to attribute model decisions…
Large-scale foundation models for scientific machine learning adapt to physical settings unseen during training, such as zero-shot transfer between turbulent scales. This phenomenon, in-context learning, challenges conventional…
Stochastic systems often exhibit multiple viable metastable states that are long-lived. Over very long timescales, fluctuations may push the system to transition between them, drastically changing its macroscopic configuration. In realistic…
Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of…
The self-attention mechanism prevails in modern machine learning. It has an interesting functionality of adaptively selecting tokens from an input sequence by modulating the degree of attention localization, which many researchers speculate…
The structure, thermodynamics and slow activated dynamics of the equilibrated metastable regime of glass-forming fluids remains a poorly understood problem of high theoretical and experimental interest. We apply a highly accurate…
Active particles self-propel themselves with a stochastically evolving velocity, generating a persistent motion leading to a non-diffusive behavior of the position distribution. Nevertheless, an effective diffusive behavior emerges at times…
With a toppling rule which generates metastable sites, we explore the properties of a gradient-driven sandpile that is minimally perturbed at one boundary. In two dimensions we find that the transport of grains takes place along deep…
We consider the dynamics of a periodic chain of N coupled overdamped particles under the influence of noise, in the limit of large N. Each particle is subjected to a bistable local potential, to a linear coupling with its nearest…
We take a deep look into the behavior of self-attention heads in the transformer architecture. In light of recent work discouraging the use of attention distributions for explaining a model's behavior, we show that attention distributions…
Decoding human activity accurately from wearable sensors can aid in applications related to healthcare and context awareness. The present approaches in this domain use recurrent and/or convolutional models to capture the spatio-temporal…
In order to successfully perform tasks specified by natural language instructions, an artificial agent operating in a visual world needs to map words, concepts, and actions from the instruction to visual elements in its environment. This…
The Transformer architecture has revolutionized deep learning, delivering the state-of-the-art performance in areas such as natural language processing, computer vision, and time series prediction. However, its core component,…
The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits…
Transformer-based models have emerged as a leading architecture for natural language processing, natural language generation, and image generation tasks. A fundamental element of the transformer architecture is self-attention, which allows…
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This…
Fixed-time stable dynamical systems are capable of achieving exact convergence to an equilibrium point within a fixed time that is independent of the initial conditions of the system. This property makes them highly appealing for designing…
Metastability, characterized by a variability of regimes in time, is a ubiquitous type of neural dynamics. It has been formulated in many different ways in the neuroscience literature, however, which may cause some confusion. In this…
In this communication we analyze the behavior of excited drops contained in spherical volumes. We study different properties of the dynamical systems i.e. the maximum Lyapunov exponent MLE, the asymptotic distance in momentum space…
In this paper, we introduce a data-driven modeling approach for dynamics problems with latent variables. The state-space of the proposed model includes artificial latent variables, in addition to observed variables that can be fitted to a…