中文
相关论文

相关论文: Dynamic metastability in the self-attention model

200 篇论文

Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting…

机器学习 · 计算机科学 2024-02-14 Borjan Geshkovski , Cyril Letrouit , Yury Polyanskiy , Philippe Rigollet

Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…

机器学习 · 计算机科学 2026-05-11 Ayan Pendharkar

We develop a mathematical framework that interprets Transformer attention as an interacting particle system and studies its continuum (mean-field) limits. By idealizing attention on the sphere, we connect Transformer dynamics to Wasserstein…

机器学习 · 计算机科学 2026-02-02 Philippe Rigollet

We investigate a model of interacting clusters which compete for growth. For a finite assembly of coupled clusters, the largest one always wins, so that all but this one die out in a finite time. This scenario of `survival of the biggest'…

统计力学 · 物理学 2007-05-23 J. M. Luck , Anita Mehta

Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers with hardmax self-attention and normalization sublayers as…

计算与语言 · 计算机科学 2026-05-14 Albert Alcalde , Giovanni Fantuzzi , Enrique Zuazua

Transformers owe much of their empirical success in natural language processing to the self-attention blocks. Recent perspectives interpret attention blocks as interacting particle systems, whose mean-field limits correspond to gradient…

机器学习 · 计算机科学 2026-03-18 Viktor Stein , Wuchen Li , Gabriele Steidl

Learning reduced descriptions of chaotic many-body dynamics is fundamentally challenging: although microscopic equations are Markovian, collective observables exhibit strong memory and exponential sensitivity to initial conditions and…

计算物理 · 物理学 2026-01-28 Ho Jang , Gia-Wei Chern

Transformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite this empirical success, the theoretical foundations of…

机器学习 · 计算机科学 2026-04-14 Zhen Qin , Jinxin Zhou , Jiachen Jiang , Zhihui Zhu

In this article, we perform quantitative analyses of metastable behavior of an interacting particle system known as the inclusion process. For inclusion processes, it is widely believed that the system nucleates the condensation of…

概率论 · 数学 2021-02-24 Seonwoo Kim , Insuk Seo

The reformulation of the mode-coupling theory (MCT) of the liquid-glass transition which incorporates the element of metastability is applied to the hard-sphere system. It is shown that the glass transition in this system is not a sharp one…

凝聚态物理 · 物理学 2016-08-31 Joonhyun Yeo

In machine learning, a self-attention dynamics is a continuous-time multiagent-like model of the attention mechanisms of transformers. In this paper we show that such dynamics is related to a multiagent version of the Oja flow, a dynamical…

机器学习 · 计算机科学 2025-11-17 Claudio Altafini

Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that a very lightweight…

计算与语言 · 计算机科学 2019-02-26 Felix Wu , Angela Fan , Alexei Baevski , Yann N. Dauphin , Michael Auli

We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth as a time variable, the residual stream defines a…

概率论 · 数学 2026-04-03 Hugo Koubbi , Borjan Geshkovski , Philippe Rigollet

In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretical point of view, the mathematical models proposed in the…

Transformers are one of the most successful architectures of modern neural networks. At their core there is the so-called attention mechanism, which recently interested the physics community as it can be written as the derivative of an…

机器学习 · 计算机科学 2024-09-25 Francesco D'Amico , Matteo Negri

Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens. This representation is then exploited by the attention function, which learns dependencies between tokens and…

机器学习 · 计算机科学 2025-01-31 Valérie Castin , Pierre Ablin , José Antonio Carrillo , Gabriel Peyré

An extremely broad and important class of phenomena in nature involves the settling and aggregation of matter under gravitation in fluid systems. Some examples include: sedimenting marine snow particles in lakes and oceans (central to…

流体动力学 · 物理学 2019-05-21 Roberto Camassa , Daniel M. Harris , Robert Hunt , Zeliha Kilic , Richard M. McLaughlin

Dynamical clustering represents a characteristic feature of active matter consisting of self-propelled agents that convert energy from the environment into mechanical motion. At the micron scale, typical of overdamped dynamics, particles…

软凝聚态物质 · 物理学 2024-06-04 Lorenzo Caprini , Davide Breoni , Anton Ldov , Christian Scholz , Hartmut Löwen

This work presents a modification of the self-attention dynamics proposed by Geshkovski et al. (arXiv:2312.10794) to better reflect the practically relevant, causally masked attention used in transformer architectures for generative AI.…

机器学习 · 计算机科学 2024-11-12 Nikita Karagodin , Yury Polyanskiy , Philippe Rigollet

Self-attention is a method of encoding sequences of vectors by relating these vectors to each-other based on pairwise similarities. These models have recently shown promising results for modeling discrete sequences, but they are non-trivial…

计算与语言 · 计算机科学 2018-06-19 Matthias Sperber , Jan Niehues , Graham Neubig , Sebastian Stüker , Alex Waibel
‹ 上一页 1 2 3 10 下一页 ›