Related papers: The Diffusion-Attention Connection
The dynamics of interacting particles in orbital magnetic fields are notoriously difficult to study, as this physics is inherently connected to electronic correlations in two-dimensional systems, for which no straightforward theoretical…
Riemannian diffusion models draw inspiration from standard Euclidean space diffusion models to learn distributions on general manifolds. Unfortunately, the additional geometric complexity renders the diffusion transition term inexpressible…
The strongly interacting matter created in relativistic heavy-ion collisions possesses several conserved quantum numbers, such as baryon number, strangeness, and electric charge. The diffusion process of these charges can be characterized…
We explain how to use diffusion models to learn inverse renormalization group flows of statistical and quantum field theories. Diffusion models are a class of machine learning models which have been used to generate samples from complex…
Diffusion magnetic resonance imaging (dMRI) is a relatively modern technique used to study tissue microstructure in a non-invasive way. Non-Gaussian diffusion representation is related to the restricted diffusion and can provide information…
We study dynamics of a generic quadratic diffeomorphism, a 3D generalization of the planar H\'{e}non map. Focusing on the dissipative, orientation preserving case, we give a comprehensive parameter study of codimension-one and two…
In recent years, diffusion models have become the leading approach for distribution learning. This paper focuses on structure-preserving diffusion models (SPDM), a specific subset of diffusion processes tailored for distributions with…
Group equivariant neural networks are used as building blocks of group invariant neural networks, which have been shown to improve generalisation performance and data efficiency through principled parameter sharing. Such works have mostly…
In both Computer Vision and the wider Deep Learning field, the Transformer architecture is well-established as state-of-the-art for many applications. For Multitask Learning, however, where there may be many more queries necessary compared…
Uncertainty calibration in pre-trained transformers is critical for their reliable deployment in risk-sensitive applications. Yet, most existing pre-trained transformers do not have a principled mechanism for uncertainty propagation through…
We argue that Transformers are essentially graph-to-graph models, with sequences just being a special case. Attention weights are functionally equivalent to graph edges. Our Graph-to-Graph Transformer architecture makes this ability…
We discuss application of methods from the Kraichnan model of turbulent advection to the study of non-equilibrium concentration fluctuations arising during diffusion in liquid mixtures at high Schmidt numbers. This approach treats nonlinear…
Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interaction pattern induces strong synchronization dynamics that…
Data-driven learning of partial differential equations' solution operators has recently emerged as a promising paradigm for approximating the underlying solutions. The solution operators are usually parameterized by deep learning models…
Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of locality-inducing modules in hybrid architectures has…
The Transformer, with its scaled dot-product attention mechanism, has become a foundational architecture in modern AI. However, this mechanism is computationally intensive and incurs substantial energy costs. We propose a new Transformer…
Diffusion Transformers (DiT) have become a leading architecture in image generation. However, the quadratic complexity of attention mechanisms, which are responsible for modeling token-wise relationships, results in significant latency when…
Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of…
We develop a consistent quantum description of surface plasmons interacting with quantum emitters and external electromagnetic field. Within the framework of macroscopic electrodynamics in dispersive and absorptive medium, we derive, in the…
We propose a simple modification to the conventional attention mechanism applied by Transformers: Instead of quantifying pairwise query-key similarity with scaled dot-products, we quantify it with the logarithms of scaled dot-products of…