Related papers: The Diffusion-Attention Connection
We show that the relation between the Schr\"odinger equation and diffusion processes has an algebraic nature and can be revealed via the structure of "duplex numbers." This helps one to clarify that quantum mechanics cannot be reduced to…
Transformers are widely used in natural language processing, where they consistently achieve state-of-the-art performance. This is mainly due to their attention-based architecture, which allows them to model rich linguistic relations…
The master equation describing non-equilibrium one-dimensional problems like diffusion limited reactions or critical dynamics of classical spin systems can be written as a Schr\"odinger equation in which the wave function is the probability…
Subdiffusion on graphs is often modeled by time-fractional diffusion equations, yet its structural and dynamical consequences remain unclear. We show that subdiffusive transport on graphs is a memory-driven process generated by a random…
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This…
We introduce a class of partial differential equations on metric graphs associated with mixed evolution: on some edges we consider diffusion processes, on other ones transport phenomena. This yields a system of equations with possibly…
We present new connections among anomalous diffusion (AD), normal diffusion (ND) and the Central Limit Theorem. This is done by defining a point transformation to a new position variable, which we postulate to be Cartesian, motivated by…
We construct diffusions with values in the nonnegative orthant, normal reflection along each of the axes, and two pairs of local drift/variance characteristics assigned according to rank; one of the variances is allowed to vanish, but not…
Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This…
Seismic data interpolation is a critical pre-processing step for improving seismic imaging quality and remains a focus of academic innovation. To address the computational inefficiencies caused by extensive iterative resampling in current…
Transformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite this empirical success, the theoretical foundations of…
Transformer-based architectures have become the dominant paradigm for Continuous-Time Dynamic Graph (CTDG) learning, yet their performance remains limited on temporally shifted datasets. In this work, we identify attention dispersion as a…
Multi-head attention empowers the recent success of transformers, the state-of-the-art models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot products between the…
Following Le Jan and Watanabe we define a connection associated with a non-degenrrate diffusion operators. This connection is characterized here and shown to be the Levi-Civita connection for gradient systems. This both explains why such…
We consider in this paper a diffusion-convection reaction equation in one space dimension. The main assumptions are about the reaction term, which is monostable, and the diffusivity, which changes sign once or twice; then, we deal with a…
We assume that a symplectic real-analytic map has an invariant normally hyperbolic cylinder and an associated transverse homoclinic cylinder. It is well known that such cylinder is preserved under small perturbations. We prove that for a…
We propose Joint MLP/Attention (JoMA) dynamics, a novel mathematical framework to understand the training procedure of multilayer Transformer architectures. This is achieved by integrating out the self-attention layer in Transformers,…
With the development of the self-attention mechanism, the Transformer model has demonstrated its outstanding performance in the computer vision domain. However, the massive computation brought from the full attention mechanism became a…
The unified description of diffusion processes that cross over from a ballistic behavior at short times to normal or anomalous diffusion (sub- or superdiffusion) at longer times is constructed on the basis of a non-Markovian generalization…
We introduce the concept of Hypoelliptic Diffusion Maps (HDM), a framework generalizing Diffusion Maps in the context of manifold learning and dimensionality reduction. Standard non-linear dimensionality reduction methods (e.g., LLE,…