Related papers: Dynamic metastability in the self-attention model
We study the behaviour of interacting self-propelled particles, whose self-propulsion speed decreases with their local density. By combining direct simulations of the microscopic model with an analysis of the hydrodynamic equations obtained…
We study a coupled dynamics of a network and a particle system. Particles of density $\rho$ diffuse freely along edges, each of which is rewired at a rate given by a decreasing function of particle flux. We find that the coupled dynamics…
Enhancing the kinetic stability of glasses often necessitates deepening thermodynamic stability, which typically compromises ductility due to increased structural rigidity. Decoupling these properties remains a critical challenge for…
User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existing user behavior modeling solutions, attention networks are…
We introduce a simple model of active transport for an ensemble of particles driven by an external shear flow. Active refers to the fact that the flow of the particles is modified by the distribution of particles itself. The model consists…
Recent observations of hydrostatic structure and virial equilibrium in supersonically turbulent, self-gravitating molecular clouds imply a stability that contrasts with the transcience of turbulent structure. To investigate this…
Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond configuration-specific studies. Inspired by empirical…
Rare transitions between long-lived metastable states underlie a great variety of physical, chemical and biological processes. Our quantitative understanding of reactive mechanisms has been driven forward by the insights of transition state…
Lattices are an efficient and effective method to encode ambiguity of upstream systems in natural language processing tasks, for example to compactly capture multiple speech recognition hypotheses, or to represent multiple linguistic…
We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and show that a simple first-order optimization algorithm can converge to the globally optimal self-attention…
Decomposing knowledge into interchangeable pieces promises a generalization advantage when there are changes in distribution. A learning agent interacting with its environment is likely to be faced with situations requiring novel…
Understanding the training dynamics of transformers is important to explain the impressive capabilities behind large language models. In this work, we study the dynamics of training a shallow transformer on a task of recognizing…
Transformers are emerging as the new workhorse of NLP, showing great success across tasks. Unlike LSTMs, transformers process input sequences entirely through self-attention. Previous work has suggested that the computational capabilities…
Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and…
The dynamics of complex systems generally include high-dimensional, non-stationary and non-linear behavior, all of which pose fundamental challenges to quantitative understanding. To address these difficulties we detail a new approach based…
We apply a recently developed theory for metastability in open quantum systems to a one-dimensional dissipative quantum Ising model. Earlier results suggest this model features either a non-equilibrium phase transition or a smooth but sharp…
Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model…
The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations…
Self-attention model have shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the…
Metastability and its relaxation mechanisms challenge our understanding of the stability of quantum many-body systems, revealing a gap between the microscopic dynamics of the individual components and the effective descriptions used for…