Related papers: The Diffusion-Attention Connection
We propose a graph-oriented attention-based explainability method for tabular data. Tasks involving tabular data have been solved mostly using traditional tree-based machine learning models which have the challenges of feature selection and…
The study of diffusion in Hamiltonian systems has been a problem of interest for a number of years. In this paper we explore the influence of self-consistency on the diffusion properties of systems described by coupled symplectic maps.…
Attention regulates information transfer between tokens. For this, query and key vectors are compared, typically in terms of a scalar product, $\mathbf{Q}^T\mathbf{K}$, together with a subsequent softmax normalization. In geometric terms,…
Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention…
Pendry and MacKinnon meaningful discretization of Maxwell's equations was put forward specifically as part of a finite-element numerical algorithm. By contrast with a numerical approach, in the same spirit evoked by the relationships…
We study the diffusion phenomena on the negatively curved surface made up of congruent heptagons. Unlike the usual two-dimensional plane, this structure makes the boundary increase exponentially with the distance from the center, and hence…
This paper presents an artistic and technical investigation into the attention mechanisms of video diffusion transformers. Inspired by early video artists who manipulated analog video signals to create new visual aesthetics, this study…
We note that building a magnetic Laplacian from the Markov transition matrix, rather than the graph adjacency matrix, yields several benefits for the magnetic eigenmaps algorithm. The two largest benefits are that the embedding becomes more…
Pairwise dot product-based attention allows Transformers to exchange information between tokens in an input-dependent way, and is key to their success across diverse applications in language and vision. However, a typical Transformer model…
Real-world data generation often involves complex inter-dependencies among instances, violating the IID-data hypothesis of standard learning paradigms and posing a challenge for uncovering the geometric structures for learning desired…
Transformer architecture has become ubiquitous in the natural language processing field. To interpret the Transformer-based models, their attention patterns have been extensively analyzed. However, the Transformer architecture is not only…
A thorough review of the q-space technique is presented starting from a discussion of Fick's laws. The work presented here is primarily conceptual, theoretical and hopefully pedagogical. We offered the notion of molecular concentration to…
Multi-head attention is a driving force behind state-of-the-art transformers, which achieve remarkable performance across a variety of natural language processing (NLP) and computer vision tasks. It has been observed that for many…
Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation…
The modelling of linear and nonlinear reaction-subdiffusion processes is more subtle than normal diffusion and causes different phenomena. The resulting equations feature a spatial Laplacian with a temporal memory term through a time…
The dot product attention mechanism, originally designed for natural language processing tasks, is a cornerstone of modern Transformers. It adeptly captures semantic relationships between word pairs in sentences by computing a similarity…
Spatial diffusion of particles in periodic potential models has provided a good framework for studying the role of chaos in global properties of classical systems. Here a bidimensional "soft" billiard, classically modeled from an optical…
Transformers have revolutionized deep learning in numerous fields, including natural language processing, computer vision, and audio processing. Their strength lies in their attention mechanism, which allows for the discovering of complex…
Attention mechanisms have become a popular component in deep neural networks, yet there has been little examination of how different influencing factors and methods for computing attention from these factors affect performance. Toward a…
A few models have tried to tackle the link prediction problem, also known as knowledge graph completion, by embedding knowledge graphs in comparably lower dimensions. However, the state-of-the-art results are attained at the cost of…