中文
相关论文

相关论文: The Diffusion-Attention Connection

200 篇论文

The dynamics of interacting particles in orbital magnetic fields are notoriously difficult to study, as this physics is inherently connected to electronic correlations in two-dimensional systems, for which no straightforward theoretical…

量子气体 · 物理学 2026-03-27 Łukasz Iwanek , Marcin Mierzejewski , Adam S. Sajna

Riemannian diffusion models draw inspiration from standard Euclidean space diffusion models to learn distributions on general manifolds. Unfortunately, the additional geometric complexity renders the diffusion transition term inexpressible…

机器学习 · 计算机科学 2023-11-01 Aaron Lou , Minkai Xu , Stefano Ermon

The strongly interacting matter created in relativistic heavy-ion collisions possesses several conserved quantum numbers, such as baryon number, strangeness, and electric charge. The diffusion process of these charges can be characterized…

高能物理 - 唯象学 · 物理学 2024-11-06 Sourav Dey , Amaresh Jaiswal , Hiranmaya Mishra

We explain how to use diffusion models to learn inverse renormalization group flows of statistical and quantum field theories. Diffusion models are a class of machine learning models which have been used to generate samples from complex…

高能物理 - 理论 · 物理学 2023-09-07 Jordan Cotler , Semon Rezchikov

Diffusion magnetic resonance imaging (dMRI) is a relatively modern technique used to study tissue microstructure in a non-invasive way. Non-Gaussian diffusion representation is related to the restricted diffusion and can provide information…

信号处理 · 电气工程与系统科学 2020-09-17 Tomasz Pieciak , Maryam Afzali , Fabian Bogusz , Aja-Fernández , Derek K. Jones

We study dynamics of a generic quadratic diffeomorphism, a 3D generalization of the planar H\'{e}non map. Focusing on the dissipative, orientation preserving case, we give a comprehensive parameter study of codimension-one and two…

混沌动力学 · 物理学 2023-06-08 Amanda E Hampton , James D Meiss

In recent years, diffusion models have become the leading approach for distribution learning. This paper focuses on structure-preserving diffusion models (SPDM), a specific subset of diffusion processes tailored for distributions with…

机器学习 · 计算机科学 2025-03-12 Haoye Lu , Spencer Szabados , Yaoliang Yu

Group equivariant neural networks are used as building blocks of group invariant neural networks, which have been shown to improve generalisation performance and data efficiency through principled parameter sharing. Such works have mostly…

机器学习 · 计算机科学 2021-06-17 Michael Hutchinson , Charline Le Lan , Sheheryar Zaidi , Emilien Dupont , Yee Whye Teh , Hyunjik Kim

In both Computer Vision and the wider Deep Learning field, the Transformer architecture is well-established as state-of-the-art for many applications. For Multitask Learning, however, where there may be many more queries necessary compared…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Christian Bohn , Thomas Kurbiel , Klaus Friedrichs , Hasan Tercan , Tobias Meisen

Uncertainty calibration in pre-trained transformers is critical for their reliable deployment in risk-sensitive applications. Yet, most existing pre-trained transformers do not have a principled mechanism for uncertainty propagation through…

We argue that Transformers are essentially graph-to-graph models, with sequences just being a special case. Attention weights are functionally equivalent to graph edges. Our Graph-to-Graph Transformer architecture makes this ability…

计算与语言 · 计算机科学 2023-10-30 James Henderson , Alireza Mohammadshahi , Andrei C. Coman , Lesly Miculicich

We discuss application of methods from the Kraichnan model of turbulent advection to the study of non-equilibrium concentration fluctuations arising during diffusion in liquid mixtures at high Schmidt numbers. This approach treats nonlinear…

统计力学 · 物理学 2022-10-18 Gregory Eyink , Amir Jafari

Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interaction pattern induces strong synchronization dynamics that…

机器学习 · 计算机科学 2026-05-26 Jingkun Liu , Yisong Yue , Max Welling , Yue Song

Data-driven learning of partial differential equations' solution operators has recently emerged as a promising paradigm for approximating the underlying solutions. The solution operators are usually parameterized by deep learning models…

机器学习 · 计算机科学 2023-05-01 Zijie Li , Kazem Meidani , Amir Barati Farimani

Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of locality-inducing modules in hybrid architectures has…

机器学习 · 计算机科学 2026-02-17 Yukun Zhang , Xueqing Zhou

The Transformer, with its scaled dot-product attention mechanism, has become a foundational architecture in modern AI. However, this mechanism is computationally intensive and incurs substantial energy costs. We propose a new Transformer…

机器学习 · 计算机科学 2025-08-07 Xin Gao , Xingming Xu , Shirin Amiraslani , Hong Xu

Diffusion Transformers (DiT) have become a leading architecture in image generation. However, the quadratic complexity of attention mechanisms, which are responsible for modeling token-wise relationships, results in significant latency when…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Songhua Liu , Zhenxiong Tan , Xinchao Wang

Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics of a single layer of…

机器学习 · 统计学 2024-06-04 Juno Kim , Taiji Suzuki

We develop a consistent quantum description of surface plasmons interacting with quantum emitters and external electromagnetic field. Within the framework of macroscopic electrodynamics in dispersive and absorptive medium, we derive, in the…

介观与纳米尺度物理 · 物理学 2021-01-21 Tigran V. Shahbazyan

We propose a simple modification to the conventional attention mechanism applied by Transformers: Instead of quantifying pairwise query-key similarity with scaled dot-products, we quantify it with the logarithms of scaled dot-products of…

机器学习 · 计算机科学 2024-04-30 Franz A. Heinsen