中文
相关论文

相关论文: The Mean-Field Dynamics of Transformers

200 篇论文

Transformers are increasingly adopted for modeling and forecasting time-series, yet their internal mechanisms remain poorly understood from a dynamical systems perspective. In contrast to classical autoregressive and state-space models,…

机器学习 · 计算机科学 2025-12-25 Gregory Duthé , Nikolaos Evangelou , Wei Liu , Ioannis G. Kevrekidis , Eleni Chatzi

Despite the Transformer's dominance across machine learning, its architecture remains largely heuristic and lacks a unified theoretical foundation. We introduce Score-based Variational Flow (SVFlow), a continuous-time dynamical system for…

机器学习 · 计算机科学 2026-04-28 Huadong Liao

To overcome the quadratic cost of self-attention, recent works have proposed various sparse attention modules, most of which fall under one of two groups: 1) sparse attention under a hand-crafted patterns and 2) full attention followed by a…

机器学习 · 计算机科学 2022-10-28 Sungjun Cho , Seonwoo Min , Jinwoo Kim , Moontae Lee , Honglak Lee , Seunghoon Hong

We develop an operator-based framework to coarse-grain interacting particle systems that exhibit clustering dynamics. Starting from the particle-based transfer operator, we first construct a sequence of reduced representations: the operator…

Transformer models systematically favor certain token positions, yet the architectural origins of this position bias remain poorly understood. This bias is closely connected to the Lost-in-the-Middle phenomenon, where models underutilize…

机器学习 · 计算机科学 2026-05-28 Hanna Herasimchyk , Robin Labryga , Tomislav Prusina , Sören Laue

The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a softmax operation applied to a scaled dot product between query…

机器学习 · 计算机科学 2026-04-02 Hariprasath Govindarajan , Per Sidén , Jacob Roll , Fredrik Lindsten

Simulations of extended quantum systems are typically performed by extrapolating results of a sequence of finite-system-size simulations to the thermodynamic limit. In the quantum Monte Carlo community, twist-averaging was pioneered as an…

强关联电子 · 物理学 2022-12-14 Sergei Iskakov , Hanna Terletska , Emanuel Gull

We consider the gelation of particles which are permanently connected by random crosslinks, drawn from an ensemble of finite-dimensional continuum percolation. To average over the randomness, we apply the replica trick, and interpret the…

软凝聚态物质 · 物理学 2009-11-07 Kurt Broderix , Martin Weigt , Annette Zippelius

In this work, we present a generalized formulation of the Transformer algorithm by reinterpreting its core mechanisms within the framework of Path Integral formalism. In this perspective, the attention mechanism is recast as a process that…

高能物理 - 唯象学 · 物理学 2025-05-02 Won-Gi Paeng , Daesuk Kwon , Kyungwon Jeong , Honggyo Suh

Quantum phase transitions in many-body systems are fundamentally characterized by complex correlation structures, which pose computational challenges for conventional methods in large systems. To address this, we propose a hybrid…

量子物理 · 物理学 2026-02-03 Jin-Long Chen , Xin Li , Zhang-Qi Yin

Rich out of equilibrium collective dynamics of strongly interacting large assemblies emerge in many areas of science. Some intriguing and not fully understood examples are the glassy arrest in atomic, molecular or colloidal systems,…

统计力学 · 物理学 2023-05-03 Leticia F. Cugliandolo

The ubiquitous occurrence of cluster patterns in nature still lacks a comprehensive understanding. It is known that the dynamics of many such natural systems is captured by ensembles of Stuart-Landau oscillators. Here, we investigate…

斑图形成与孤子 · 物理学 2019-02-13 Felix P. Kemeth , Sindre W. Haugland , Katharina Krischer

Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length…

计算与语言 · 计算机科学 2025-10-08 Alexander M. Fichtl , Jeremias Bohn , Josefin Kelber , Edoardo Mosca , Georg Groh

Transformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Hila Chefer , Shir Gur , Lior Wolf

The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Laziz U. Abdullaev , Maksim Tkachenko , Tan M. Nguyen

We consider two mean-field like models which belong to the universality class of absorbing phase transitions with a conserved field. In both cases we derive analytically the order parameter as function of the control parameter and of an…

统计力学 · 物理学 2015-06-24 S. Lubeck , A. Hucht

We consider mean-field models for data--clustering problems starting from a generalization of the bounded confidence model for opinion dynamics. The microscopic model includes information on the position as well as on additional features of…

数值分析 · 数学 2020-03-16 Michael Herty , Lorenzo Pareschi , Giuseppe Visconti

Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a rigorous mathematical framework for analyzing gradient-based…

最优化与控制 · 数学 2026-05-19 Raphaël Barboni , Maarten V. de Hoop , Takashi Furuya , Gabriel Peyré

The dynamic behavior of cluster algorithms is analyzed in the classical mean field limit. Rigorous analytical results below $T_c$ establish that the dynamic exponent has the value $z_{sw}=1$ for the Swendsen-Wang algorithm and $z_{uw}=0$…

凝聚态物理 · 物理学 2009-10-28 N. Persky , R. Ben-Av , I. Kanter , E. Domany

Transformer is a powerful model for text understanding. However, it is inefficient due to its quadratic complexity to input sequence length. Although there are many methods on Transformer acceleration, they are still either inefficient on…

计算与语言 · 计算机科学 2021-09-07 Chuhan Wu , Fangzhao Wu , Tao Qi , Yongfeng Huang , Xing Xie