中文
相关论文

相关论文: Scaling Limits of Long-Context Transformers

200 篇论文

We study real-time correlation functions in scalar quantum field theories at temperature $T=1/\beta$. We show that the behaviour of soft, long wavelength modes is determined by classical statistical field theory. The loss of quantum…

高能物理 - 理论 · 物理学 2014-11-18 W. Buchmuller , A. Jakovac

The $\beta$ ensembles are a class of eigenvalue probability densities which generalise the invariant ensembles of classical random matrix theory. In the case of the Gaussian and Laguerre weights, the corresponding eigenvalue densities are…

数学物理 · 物理学 2018-12-20 Peter J. Forrester , Allan K. Trinh

While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or…

计算与语言 · 计算机科学 2023-11-21 Amirkeivan Mohtashami , Martin Jaggi

We investigate the complex-temperature singularities of the susceptibility of the 2D Ising model on a square lattice. From an analysis of low-temperature series expansions, we find evidence that as one approaches the point $u=u_s=-1$ (where…

高能物理 - 格点 · 物理学 2009-10-22 V. Matveev , R. Shrock

Extensive simulations are made on Ising Spin Glasses (ISG) with Gaussian, Laplacian and bimodal interaction distributions in dimension four. Standard finite size scaling analyses near and at criticality provide estimates of the critical…

无序系统与神经网络 · 物理学 2014-08-06 P. H. Lundow , I. A. Campbell

Large language models spend most of their inference cost on attention over long contexts, yet empirical behavior suggests that only a small subset of tokens meaningfully contributes to each query. We formalize this phenomenon by modeling…

人工智能 · 计算机科学 2026-02-17 Vashista Nobaub

This paper studies beta ensembles on the real line in a high temperature regime, that is, the regime where $\beta N \to const \in (0, \infty)$, with $N$ the system size and $\beta$ the inverse temperature. In this regime, the convergence to…

概率论 · 数学 2020-04-17 Fumihiko Nakano , Khanh Duy Trinh

We study transformer language models, analyzing attention heads whose attention patterns are spread out, and whose attention scores depend weakly on content. We argue that the softmax denominators of these heads are stable when the…

计算与语言 · 计算机科学 2025-10-07 Alex Gibson

The talk presented at ICMP 97 focused on the scaling limits of critical percolation models, and some other systems whose salient features can be described by collections of random lines. In the scaling limit we keep track of features seen…

数学物理 · 物理学 2007-05-23 Michael Aizenman

The five-dimensional Ising model with free boundary conditions has recently received a renewed interest in a debate concerning the finite-size scaling of the susceptibility near the critical temperature. We provide evidence in favour of the…

统计力学 · 物理学 2016-09-13 P. H. Lundow , K. Markström

The transverse-field Ising model is widely studied as one of the simplest quantum spin systems. It is known that this model exhibits a phase transition at the critical inverse temperature $\beta_{\mathrm{c}}$, which is determined by the…

数学物理 · 物理学 2025-09-01 Yoshinori Kamijima , Akira Sakai

Scaling laws enable the optimal selection of data amount and language model size, yet the impact of the data unit, the token, on this relationship remains underexplored. In this work, we systematically investigate how the information…

Consider the classical $(2+1)$-dimensional Solid-On-Solid model above a hard wall on an $L\times L$ box of $\bbZ^2$. The model describes a crystal surface by assigning a non-negative integer height $\eta_x$ to each site $x$ in the box and 0…

To improve the robustness of transformer neural networks used for temporal-dynamics prediction of chaotic systems, we propose a novel attention mechanism called easy attention which we demonstrate in time-series reconstruction and…

Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probability measures, treating a context as a distribution over…

机器学习 · 统计学 2026-02-03 Ryotaro Kawata , Taiji Suzuki

At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on the most salient information. However, the softmax structure…

机器学习 · 计算机科学 2025-05-27 Fanqi Yan , Huy Nguyen , Pedram Akbarian , Nhat Ho , Alessandro Rinaldo

Temporal relaxation of density fluctuations in supercooled liquids near the glass transition occurs in multiple steps. The short-time $\beta$-relaxation is generally attributed to spatially local processes involving the rattling motion of a…

统计力学 · 物理学 2016-03-02 Smarajit Karmakar , Chandan Dasgupta , Srikanth Sastry

VRAM requirements for transformer models scale quadratically with context length due to the self-attention mechanism. In this paper we modify the decoder-only transformer, replacing self-attention with InAttention, which scales linearly…

机器学习 · 计算机科学 2024-10-10 Joseph Eisner

Understanding whether attention mechanisms converge to the kernel regime is foundational to the validity of influence functions for transformer accountability. Exact NTK characterization of softmax attention is precluded by its exponential…

机器学习 · 计算机科学 2026-05-08 Jose Marie Antonio Miñoza , Paulo Mario P. Medina , Sebastian C. Ibañez

In long-range percolation on $\mathbb{Z}^d$, points $x$ and $y$ are connected by an edge with probability $1-\exp(-\beta\|x-y\|^{-d-\alpha})$, where $\alpha>0$ is fixed and $\beta \geq 0$ is a parameter. As $d$ and $\alpha$ vary, the model…

概率论 · 数学 2025-08-27 Tom Hutchcroft