English
Related papers

Related papers: Analysis of mean-field models arising from self-at…

200 papers

We present an approach to modifying Transformer architectures by integrating graph-aware relational reasoning into the attention mechanism, merging concepts from graph neural networks and language modeling. Building on the inherent…

Machine Learning · Computer Science 2025-03-06 Markus J. Buehler

Transformers often appear to perform Bayesian reasoning in context, but verifying this rigorously has been impossible: natural data lack analytic posteriors, and large models conflate reasoning with memorization. We address this by…

Machine Learning · Computer Science 2026-05-19 Naman Agarwal , Siddhartha R. Dalal , Vishal Misra

We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is…

Machine Learning · Computer Science 2024-02-15 Clayton Sanford , Daniel Hsu , Matus Telgarsky

Accurate and physically consistent modeling of Earth system dynamics requires machine-learning architectures that operate directly on continuous geophysical fields and preserve their underlying geometric structure. Here we introduce…

Machine Learning · Computer Science 2025-12-24 Maximilian Witte , Johannes Meuer , Étienne Plésiat , Christopher Kadow

Recent years have seen an increased interest in the application of methods and techniques commonly associated with machine learning and artificial intelligence to spatial statistics. Here, in a celebration of the ten-year anniversary of the…

Methodology · Statistics 2022-01-25 Tin Lok James Ng , Andrew Zammit-Mangion

We investigate active electrolytes within the mean-field level of description. The focus is on how the double-layer structure of passive, thermalized charges is affected by active dynamics of all constituting ions. One feature of active…

Soft Condensed Matter · Physics 2018-05-23 Derek Frydel , Rudolf Podgornik

The aggregation equation arises naturally in kinetic theory in the study of granular media, and its interpretation as a 2-Wasserstein gradient flow for the nonlocal interaction energy is well-known. Starting from the spatially homogeneous…

Analysis of PDEs · Mathematics 2024-12-24 A. Esposito , R. S. Gvalani , A. Schlichting , M. Schmidtchen

The self-attention mechanism, now central to deep learning architectures such as Transformers, is a modern instance of a more general computational principle: learning and using pairwise affinity matrices to control how information flows…

Machine Learning · Computer Science 2025-07-29 Giorgio Roffo

Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting…

Machine Learning · Computer Science 2024-02-14 Borjan Geshkovski , Cyril Letrouit , Yury Polyanskiy , Philippe Rigollet

In stellar interiors shear flows play an important role in many physical processes. So far helioseismology provides only large-scale measurements, and so the small-scale dynamics remains insufficiently understood. To draw a connection…

Solar and Stellar Astrophysics · Physics 2016-12-01 V. Witzke , L. J. Silvers

Transformers architecture apply self-attention to tokens represented as vectors, before a fully connected (neuronal network) layer. These two parts can be layered many times. Traditionally, self-attention is seen as a mechanism for…

Computation and Language · Computer Science 2025-01-22 Evgeniy Shin , Heinrich Matzinger

We present a systematic derivation of the gradient flows associated to a broad class of interfacial energies, emphasizing the relation between intrinsic and extrinsic variations of the interface. We show that the intrinsic variables…

Analysis of PDEs · Mathematics 2025-01-28 Vinh Nguyen , Keith Promislow , Brian Wetton

Dynamic graph learning plays a pivotal role in modeling evolving relationships over time, especially for temporal link prediction tasks in domains such as traffic systems, social networks, and recommendation platforms. While…

Machine Learning · Computer Science 2025-11-18 Tao Zou , Chengfeng Wu , Tianxi Liao , Junchen Ye , Bowen Du

In this work, we provide a theoretical understanding of the framelet-based graph neural networks through the perspective of energy gradient flow. By viewing the framelet-based models as discretized gradient flows of some energy, we show it…

Machine Learning · Computer Science 2022-10-11 Andi Han , Dai Shi , Zhiqi Shao , Junbin Gao

Transformer-based models have emerged as a leading architecture for natural language processing, natural language generation, and image generation tasks. A fundamental element of the transformer architecture is self-attention, which allows…

Machine Learning · Computer Science 2025-07-01 Venmugil Elango

Self-attention, as the key block of transformers, is a powerful mechanism for extracting features from the inputs. In essence, what self-attention does is to infer the pairwise relations between the elements of the inputs, and modify the…

Machine Learning · Computer Science 2021-03-09 Lemeng Wu , Xingchao Liu , Qiang Liu

The modeling of high-dimensional spatio-temporal processes presents a fundamental dichotomy between the probabilistic rigor of classical geostatistics and the flexible, high-capacity representations of deep learning. While Gaussian…

Machine Learning · Computer Science 2025-12-22 Yuri Calleo

Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus-dilution cycle in attention learning and…

Machine Learning · Computer Science 2026-05-05 Zheng-An Chen , Pengxiao Lin , Zhi-Qin John Xu , Tao Luo

Linear attention mechanisms have emerged as efficient alternatives to full self-attention in Graph Transformers, offering linear time complexity. However, existing linear attention models often suffer from a significant drop in…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Zhaolin Hu , Kun Li , Hehe Fan , Yi Yang

While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between attention mechanisms and multinomial regression. Specifically,…

Machine Learning · Computer Science 2025-10-28 Jonas A. Actor , Anthony Gruber , Eric C. Cyr