Related papers: Dynamic metastability in the self-attention model
We perform extensive numerical simulations of a paradigmatic model glass former, the hard-sphere fluid with 10% polydispersity. We sample from the ensemble of trajectories with fixed observation time, whereby single trajectories are…
Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization…
We show that a scaling approach successfully characterizes clustering and intermittency in space and time, in systems of noninteracting particles driven by fluctuating surfaces. We study both the steady state and the approach to it, for…
We study Tao's finitary viewpoint of convergence in metric spaces, as captured by the notion of metastability. We adopt the perspective of continuous model theory. We show that, in essence, metastable convergence with a given rate is the…
Graph-based next-step prediction models have recently been very successful in modeling complex high-dimensional physical systems on irregular meshes. However, due to their short temporal attention span, these models suffer from error…
We show that metastable ring-shaped clusters can be constructed from two-dimensional quantum droplets in systems described by the Gross-Pitaevskii equations augmented with Lee-Huang-Yang quantum corrections. The clusters exhibit dynamical…
In contexts where data samples represent a physically stable state, it is often assumed that the data points represent the local minima of an energy landscape. In control theory, it is well-known that energy can serve as an effective…
The self-attention mechanism, at the heart of the Transformer model, is able to effectively model pairwise interactions between tokens. However, numerous recent works have shown that it is unable to perform basic tasks involving detecting…
Modern deep learning approaches have achieved groundbreaking performance in modeling and classifying sequential data. Specifically, attention networks constitute the state-of-the-art paradigm for capturing long temporal dynamics. This paper…
In neuroscience, attention has been shown to bidirectionally interact with reinforcement learning (RL) processes. This interaction is thought to support dimensionality reduction of task representations, restricting computations to relevant…
Active matter deals with systems whose particles consume energy at the individual level in order to move. To unravel features such as the emergence of collective structures several models have been suggested, such as the on-lattice model of…
Behavioral changes in animals and humans, as a consequence of an error or a verbal instruction, can be extremely rapid. Improvement in behavioral performances are usually associated in machine learning and reinforcement learning to synaptic…
We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansions to the triangular causal dependency structure of the…
We study the glassy dynamics taking place in dense assemblies of athermal active particles that are driven solely by a nonequilibrium self-propulsion mechanism. Active forces are modeled as an Ornstein-Uhlenbeck stochastic process,…
Pretrained language models based on the transformer architecture have shown great success in NLP. Textual training data often comes from the web and is thus tagged with time-specific information, but most language models ignore this…
We derive a class of macroscopic differential equations that describe collective adaptation, starting from a discrete-time stochastic microscopic model. The behavior of each agent is a dynamic balance between adaptation that locally…
We study the solid-to-liquid transition in a two-dimensional fully periodic soft-glassy model with an imposed spatially heterogeneous stress. The model we consider consists of droplets of a dispersed phase jammed together in a continuous…
We present existence, uniqueness and continuous dependence results for some kinetic equations motivated by models for the collective behavior of large groups of individuals. Models of this kind have been recently proposed to study the…
Transformer-based models, such as BERT and GPT, have been widely adopted in natural language processing (NLP) due to their exceptional performance. However, recent studies show their vulnerability to textual adversarial attacks where the…
Transformer architectures have led to remarkable progress in many state-of-art applications. However, despite their successes, modern transformers rely on the self-attention mechanism, whose time- and space-complexity is quadratic in the…