English
Related papers

Related papers: Analysis of mean-field models arising from self-at…

200 papers

Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Hila Chefer , Shir Gur , Lior Wolf

Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz properties - which are key when it comes to analyzing…

Machine Learning · Computer Science 2024-06-05 Valérie Castin , Pierre Ablin , Gabriel Peyré

This paper introduces Generalized Attention Flow (GAF), a novel feature attribution method for Transformer-based models to address the limitations of current approaches. By extending Attention Flow and replacing attention weights with the…

Machine Learning · Computer Science 2025-02-25 Behrooz Azarkhalili , Maxwell Libbrecht

In this manuscript, we show how flow equation methods can be used to study localisation in disordered quantum systems, and particularly how to use this approach to obtain the non-equilibrium dynamical evolution of observables. We review the…

Disordered Systems and Neural Networks · Physics 2020-02-27 S. J. Thomson , M. Schiró

Transformer-based architectures are the most used architectures in many deep learning fields like Natural Language Processing, Computer Vision or Speech processing. It may encourage the direct use of Transformers in the constrained tasks,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Youness Dkhissi , Valentin Vielzeuf , Elys Allesiardo , Anthony Larcher

We propose Joint MLP/Attention (JoMA) dynamics, a novel mathematical framework to understand the training procedure of multilayer Transformer architectures. This is achieved by integrating out the self-attention layer in Transformers,…

Machine Learning · Computer Science 2024-03-18 Yuandong Tian , Yiping Wang , Zhenyu Zhang , Beidi Chen , Simon Du

Approximating a probability distribution using a set of particles is a fundamental problem in machine learning and statistics, with applications including clustering and quantization. Formally, we seek a weighted mixture of Dirac measures…

Machine Learning · Statistics 2026-04-24 Ayoub Belhadji , Daniel Sharp , Youssef Marzouk

Understanding the intricate non-convex training dynamics of softmax-based models is crucial for explaining the empirical success of transformers. In this article, we analyze the gradient flow dynamics of the value-softmax model, defined as…

Machine Learning · Computer Science 2026-03-09 Aditya Varre , Mark Rofin , Nicolas Flammarion

We study the transition to turbulence in a flat plate boundary layer by means of visibility analysis of velocity time-series extracted across the flow domain. By taking into account the mutual visibility of sampled values, visibility graphs…

Fluid Dynamics · Physics 2022-10-13 Davide Perrone , Luca Ridolfi , Stefania Scarsoglio

We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and show that a simple first-order optimization algorithm can converge to the globally optimal self-attention…

Machine Learning · Computer Science 2026-03-03 Gautam Goel , Mahdi Soltanolkotabi , Peter Bartlett

We introduce Attention Graphs, a new tool for mechanistic interpretability of Graph Neural Networks (GNNs) and Graph Transformers based on the mathematical equivalence between message passing in GNNs and the self-attention mechanism in…

Machine Learning · Computer Science 2025-02-26 Batu El , Deepro Choudhury , Pietro Liò , Chaitanya K. Joshi

The incredible success of transformers on sequence modeling tasks can be largely attributed to the self-attention mechanism, which allows information to be transferred between different parts of a sequence. Self-attention allows…

Machine Learning · Computer Science 2024-08-14 Eshaan Nichani , Alex Damian , Jason D. Lee

We introduce boundary quotients and present a framework for learning densities on manifolds that arise as boundary quotients of simpler domains. We show that this framework can be used to construct normalizing flows on quotient manifolds…

Machine Learning · Computer Science 2026-05-27 William Ghanem , Benjamin Cai

We present a framework enabling variational data assimilation for gradient flows in general metric spaces, based on the minimizing movement (or Jordan-Kinderlehrer-Otto) approximation scheme. After discussing stability properties in the…

Numerical Analysis · Mathematics 2023-01-18 Jan-F. Pietschmann , Matthias Schlottbom

The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a softmax operation applied to a scaled dot product between query…

Machine Learning · Computer Science 2026-04-02 Hariprasath Govindarajan , Per Sidén , Jacob Roll , Fredrik Lindsten

Group equivariant neural networks are used as building blocks of group invariant neural networks, which have been shown to improve generalisation performance and data efficiency through principled parameter sharing. Such works have mostly…

Machine Learning · Computer Science 2021-06-17 Michael Hutchinson , Charline Le Lan , Sheheryar Zaidi , Emilien Dupont , Yee Whye Teh , Hyunjik Kim

We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansions to the triangular causal dependency structure of the…

Analysis of PDEs · Mathematics 2026-05-12 Mitia Duerinckx , Borjan Geshkovski , Stefano Rossi

Normalizing flows are a promising tool for modeling probability distributions in physical systems. While state-of-the-art flows accurately approximate distributions and energies, applications in physics additionally require smooth energies…

Machine Learning · Statistics 2021-12-01 Jonas Köhler , Andreas Krämer , Frank Noé

We investigate dynamics of large scale and slow deformations of layered structures. Starting from the respective model equations for a non-conserved system, a conserved system and a binary fluid, we derive the interface equations which are…

Soft Condensed Matter · Physics 2009-10-30 Takao Ohta , David Jasnow

We introduce and implement a method to compute stationary states of nonlinear Schr\''odinger equations on metric graphs. Stationary states are obtained as local minimizers of the nonlinear Schr\''odinger energy at fixed mass. Our method is…

Analysis of PDEs · Mathematics 2021-06-11 Christophe Besse , Romain Duboscq , Stefan Le Coz