English
Related papers

Related papers: Scaling Limits of Long-Context Transformers

200 papers

We study real-time correlation functions in scalar quantum field theories at temperature $T=1/\beta$. We show that the behaviour of soft, long wavelength modes is determined by classical statistical field theory. The loss of quantum…

High Energy Physics - Theory · Physics 2014-11-18 W. Buchmuller , A. Jakovac

The $\beta$ ensembles are a class of eigenvalue probability densities which generalise the invariant ensembles of classical random matrix theory. In the case of the Gaussian and Laguerre weights, the corresponding eigenvalue densities are…

Mathematical Physics · Physics 2018-12-20 Peter J. Forrester , Allan K. Trinh

While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or…

Computation and Language · Computer Science 2023-11-21 Amirkeivan Mohtashami , Martin Jaggi

We investigate the complex-temperature singularities of the susceptibility of the 2D Ising model on a square lattice. From an analysis of low-temperature series expansions, we find evidence that as one approaches the point $u=u_s=-1$ (where…

High Energy Physics - Lattice · Physics 2009-10-22 V. Matveev , R. Shrock

Extensive simulations are made on Ising Spin Glasses (ISG) with Gaussian, Laplacian and bimodal interaction distributions in dimension four. Standard finite size scaling analyses near and at criticality provide estimates of the critical…

Disordered Systems and Neural Networks · Physics 2014-08-06 P. H. Lundow , I. A. Campbell

Large language models spend most of their inference cost on attention over long contexts, yet empirical behavior suggests that only a small subset of tokens meaningfully contributes to each query. We formalize this phenomenon by modeling…

Artificial Intelligence · Computer Science 2026-02-17 Vashista Nobaub

This paper studies beta ensembles on the real line in a high temperature regime, that is, the regime where $\beta N \to const \in (0, \infty)$, with $N$ the system size and $\beta$ the inverse temperature. In this regime, the convergence to…

Probability · Mathematics 2020-04-17 Fumihiko Nakano , Khanh Duy Trinh

We study transformer language models, analyzing attention heads whose attention patterns are spread out, and whose attention scores depend weakly on content. We argue that the softmax denominators of these heads are stable when the…

Computation and Language · Computer Science 2025-10-07 Alex Gibson

The talk presented at ICMP 97 focused on the scaling limits of critical percolation models, and some other systems whose salient features can be described by collections of random lines. In the scaling limit we keep track of features seen…

Mathematical Physics · Physics 2007-05-23 Michael Aizenman

The five-dimensional Ising model with free boundary conditions has recently received a renewed interest in a debate concerning the finite-size scaling of the susceptibility near the critical temperature. We provide evidence in favour of the…

Statistical Mechanics · Physics 2016-09-13 P. H. Lundow , K. Markström

The transverse-field Ising model is widely studied as one of the simplest quantum spin systems. It is known that this model exhibits a phase transition at the critical inverse temperature $\beta_{\mathrm{c}}$, which is determined by the…

Mathematical Physics · Physics 2025-09-01 Yoshinori Kamijima , Akira Sakai

Scaling laws enable the optimal selection of data amount and language model size, yet the impact of the data unit, the token, on this relationship remains underexplored. In this work, we systematically investigate how the information…

Computation and Language · Computer Science 2026-05-27 Tomasz Limisiewicz , Artidoro Pagnoni , Srini Iyer , Mike Lewis , Sachin Mehta , Alisa Liu , Margaret Li , Gargi Ghosh , Luke Zettlemoyer

Consider the classical $(2+1)$-dimensional Solid-On-Solid model above a hard wall on an $L\times L$ box of $\bbZ^2$. The model describes a crystal surface by assigning a non-negative integer height $\eta_x$ to each site $x$ in the box and 0…

Probability · Mathematics 2013-02-28 Pietro Caputo , Eyal Lubetzky , Fabio Martinelli , Allan Sly , Fabio Lucio Toninelli

To improve the robustness of transformer neural networks used for temporal-dynamics prediction of chaotic systems, we propose a novel attention mechanism called easy attention which we demonstrate in time-series reconstruction and…

Machine Learning · Computer Science 2025-06-05 Marcial Sanchis-Agudo , Yuning Wang , Roger Arnau , Luca Guastoni , Jasmin Lim , Karthik Duraisamy , Ricardo Vinuesa

Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probability measures, treating a context as a distribution over…

Machine Learning · Statistics 2026-02-03 Ryotaro Kawata , Taiji Suzuki

At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on the most salient information. However, the softmax structure…

Machine Learning · Computer Science 2025-05-27 Fanqi Yan , Huy Nguyen , Pedram Akbarian , Nhat Ho , Alessandro Rinaldo

Temporal relaxation of density fluctuations in supercooled liquids near the glass transition occurs in multiple steps. The short-time $\beta$-relaxation is generally attributed to spatially local processes involving the rattling motion of a…

Statistical Mechanics · Physics 2016-03-02 Smarajit Karmakar , Chandan Dasgupta , Srikanth Sastry

VRAM requirements for transformer models scale quadratically with context length due to the self-attention mechanism. In this paper we modify the decoder-only transformer, replacing self-attention with InAttention, which scales linearly…

Machine Learning · Computer Science 2024-10-10 Joseph Eisner

Understanding whether attention mechanisms converge to the kernel regime is foundational to the validity of influence functions for transformer accountability. Exact NTK characterization of softmax attention is precluded by its exponential…

Machine Learning · Computer Science 2026-05-08 Jose Marie Antonio Miñoza , Paulo Mario P. Medina , Sebastian C. Ibañez

In long-range percolation on $\mathbb{Z}^d$, points $x$ and $y$ are connected by an edge with probability $1-\exp(-\beta\|x-y\|^{-d-\alpha})$, where $\alpha>0$ is fixed and $\beta \geq 0$ is a parameter. As $d$ and $\alpha$ vary, the model…

Probability · Mathematics 2025-08-27 Tom Hutchcroft