Related papers: The Diffusion-Attention Connection
We introduce the $2$-simplicial Transformer, an extension of the Transformer which includes a form of higher-dimensional attention generalising the dot-product attention, and uses this attention to update entity representations with tensor…
Transformer-based models have emerged as a leading architecture for natural language processing, natural language generation, and image generation tasks. A fundamental element of the transformer architecture is self-attention, which allows…
Normalizing flows, diffusion normalizing flows and variational autoencoders are powerful generative models. This chapter provides a unified framework to handle these approaches via Markov chains. We consider stochastic normalizing flows as…
The confinement of plasmas in tokamaks and stellarators depends on magnetic field lines lying in nested toroidal surfaces. The transition near the plasma edge away from the lines lying in magnetic surfaces defines properties of divertors.…
Uncovering hidden graph structures underlying real-world data is a critical challenge with broad applications across scientific domains. Recently, transformer-based models leveraging the attention mechanism have demonstrated strong…
We propose an adiabatic-elimination formalism in the dispersive regime based on a transition-centric perturbation theory. The perturbative expansion is recast into a diagrammatic framework, while adiabatic elimination is implemented through…
We identify emergent hydrodynamics governing charge transport in Brownian random circuits with various symmetries, constraints, and ranges of interactions. This is accomplished via a mapping between the averaged dynamics and the low energy…
Transformers have achieved remarkable success across natural language processing (NLP) and computer vision (CV). However, deep transformer models often suffer from an over-smoothing issue, in which token representations converge to similar…
The standard paradigm of modeling marked point processes is by parameterizing the intensity function using an attention-based (Transformer-style) architecture. Despite the flexibility of these methods, their inference is based on the…
The large proper-time behaviour of expanding boost-invariant fluids has provided many crucial insights into quark-gluon plasma dynamics. Here we formulate and explore the late-time behaviour of nonequilibrium dynamics at the level of…
Following the major successes of self-attention and Transformers for image analysis, we investigate the use of such attention mechanisms in the context of Image Quality Assessment (IQA) and propose a novel full-reference IQA method, Vision…
In this paper we develop a hybrid version of the encounter-based approach to diffusion-mediated absorption at a reactive surface, which takes into account stochastic switching of a diffusing particle's conformational state. For simplicity,…
It is generally assumed that due to factorization of long- and short-distance dynamics perturbative QCD can be applied to exclusive hadronic reactions at large momentum transfers. Within such a perturbative approach diquarks turn out to be…
Q-conditional symmetries (nonclassical symmetries) for a general class of two-component reaction-diffusion systems with constant diffusivities are studied. Using the recently introduced notion of Q-conditional symmetries of the first type…
The Transformer architecture aggregates input information through the self-attention mechanism, but there is no clear understanding of how this information is mixed across the entire model. Additionally, recent works have demonstrated that…
This paper presents a diffusion based probabilistic interpretation of spectral clustering and dimensionality reduction algorithms that use the eigenvectors of the normalized graph Laplacian. Given the pairwise adjacency matrix of all…
We present a new method to transform an expanded class of non-selfadjoint advection-diffusion operators into self-adjoint operators. The transform is based on a combination of a point transform and Lie transform in conjunction with an…
This work introduces Cross-Attentive Modulation (CAM) tokens, which are tokens whose initial value is learned, gather information through cross-attention, and modulate the nodes and edges accordingly. These tokens are meant to improve the…
In this paper, we design distributed multi-modal localization approaches for Connected and Automated vehicles. We utilize information diffusion on graphs formed by moving vehicles, based on Adapt-then-Combine strategies combined with the…
In this paper, we describe the general framework to describe the diffusion operators associated to a positive matrix. We define the equations associated to diffusion operators and present some general properties of their state vectors. We…