Related papers: The Diffusion-Attention Connection
This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less…
We introduce Hodge Diffusion Maps, a novel manifold learning algorithm designed to analyze and extract topological information from high-dimensional data-sets. This method approximates the exterior derivative acting on differential forms,…
Transformers have achieved remarkable progress in time series forecasting, yet their reliance on deterministic dot-product attention limits their capacity to model uncertainty and nonlinear dependencies across multivariate temporal…
Transformer-based models have demonstrated exceptional performance across diverse domains, becoming the state-of-the-art solution for addressing sequential machine learning problems. Even though we have a general understanding of the…
The success of Transformer language models is widely credited to their dot-product attention mechanism, which interweaves a set of key design principles: mixing information across positions (enabling multi-token interactions),…
We show that particle transport in a uniform, quantum multi-baker map, is generically ballistic in the long time limit, for any fixed value of Planck's constant. However, for fixed times, the semi-classical limit leads to diffusion. Random…
The diffusion of a system of ferromagnetic dipoles confined in a quasi-one-dimensional parabolic trap is studied using Brownian dynamics simulations. We show that the dynamics of the system is tunable by an in-plane external homogeneous…
A new diffuse interface model for a two-phase flow of two incompressible fluids with different densities is introduced using methods from rational continuum mechanics. The model fulfills local and global dissipation inequalities and is…
Transformers have catalyzed advancements in computer vision and natural language processing (NLP) fields. However, substantial computational complexity poses limitations for their application in long-context tasks, such as high-resolution…
We study Diffusion Schr\"odinger Bridge (DSB) models in the context of dynamical astrophysical systems, specifically tackling observational inverse prediction tasks within Giant Molecular Clouds (GMCs) for star formation. We introduce the…
Transformers have recently revolutionized many domains in modern machine learning and one salient discovery is their remarkable in-context learning capability, where models can solve an unseen task by utilizing task-specific prompts without…
Light-Front quantization is one of the most promising and physical tools towards studying deep inelastic scattering on the basis of quark gluon degrees of freedom. The simplified vacuum structure (nontrivial vacuum effects can only appear…
We introduce diffusion geometry as a new framework for geometric and topological data analysis. Diffusion geometry uses the Bakry-Emery $\Gamma$-calculus of Markov diffusion operators to define objects from Riemannian geometry on a wide…
The diverse quantization phenomena in 2D condensed-matter systems, being due to a uniform perpendicular magnetic field and the geometry-created lattice symmetries, are the focuses of this book. They cover the diversified magneto-electronic…
Despite the remarkable empirical performance of Transformers, their theoretical understanding remains elusive. Here, we consider a deep multi-head self-attention network, that is closely related to Transformers yet analytically tractable.…
The structure of the diffusion regions in antiparallel magnetic reconnection is investigated by means of a theory and a Vlasov simulation. The magnetic diffusion is considered as relaxation to the frozen-in state, which depends on a…
Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is data-agnostic,…
Unlike conventional "black-box" transformers with classical self-attention mechanism, we build a lightweight and interpretable transformer-like neural net by unrolling a mixed-graph-based optimization algorithm to forecast traffic with…
Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting…
Neural networks transform data through learned representations whose geometry affects separation, contraction, and generalization. Recent work studies this geometry using discrete curvature on neighborhood graphs, suggesting Ricci-flow-like…