Related papers: Scaling Limits of Long-Context Transformers
We study real-time correlation functions in scalar quantum field theories at temperature $T=1/\beta$. We show that the behaviour of soft, long wavelength modes is determined by classical statistical field theory. The loss of quantum…
The $\beta$ ensembles are a class of eigenvalue probability densities which generalise the invariant ensembles of classical random matrix theory. In the case of the Gaussian and Laguerre weights, the corresponding eigenvalue densities are…
While Transformers have shown remarkable success in natural language processing, their attention mechanism's large memory requirements have limited their ability to handle longer contexts. Prior approaches, such as recurrent memory or…
We investigate the complex-temperature singularities of the susceptibility of the 2D Ising model on a square lattice. From an analysis of low-temperature series expansions, we find evidence that as one approaches the point $u=u_s=-1$ (where…
Extensive simulations are made on Ising Spin Glasses (ISG) with Gaussian, Laplacian and bimodal interaction distributions in dimension four. Standard finite size scaling analyses near and at criticality provide estimates of the critical…
Large language models spend most of their inference cost on attention over long contexts, yet empirical behavior suggests that only a small subset of tokens meaningfully contributes to each query. We formalize this phenomenon by modeling…
This paper studies beta ensembles on the real line in a high temperature regime, that is, the regime where $\beta N \to const \in (0, \infty)$, with $N$ the system size and $\beta$ the inverse temperature. In this regime, the convergence to…
We study transformer language models, analyzing attention heads whose attention patterns are spread out, and whose attention scores depend weakly on content. We argue that the softmax denominators of these heads are stable when the…
The talk presented at ICMP 97 focused on the scaling limits of critical percolation models, and some other systems whose salient features can be described by collections of random lines. In the scaling limit we keep track of features seen…
The five-dimensional Ising model with free boundary conditions has recently received a renewed interest in a debate concerning the finite-size scaling of the susceptibility near the critical temperature. We provide evidence in favour of the…
The transverse-field Ising model is widely studied as one of the simplest quantum spin systems. It is known that this model exhibits a phase transition at the critical inverse temperature $\beta_{\mathrm{c}}$, which is determined by the…
Scaling laws enable the optimal selection of data amount and language model size, yet the impact of the data unit, the token, on this relationship remains underexplored. In this work, we systematically investigate how the information…
Consider the classical $(2+1)$-dimensional Solid-On-Solid model above a hard wall on an $L\times L$ box of $\bbZ^2$. The model describes a crystal surface by assigning a non-negative integer height $\eta_x$ to each site $x$ in the box and 0…
To improve the robustness of transformer neural networks used for temporal-dynamics prediction of chaotic systems, we propose a novel attention mechanism called easy attention which we demonstrate in time-series reconstruction and…
Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probability measures, treating a context as a distribution over…
At the core of the popular Transformer architecture is the self-attention mechanism, which dynamically assigns softmax weights to each input token so that the model can focus on the most salient information. However, the softmax structure…
Temporal relaxation of density fluctuations in supercooled liquids near the glass transition occurs in multiple steps. The short-time $\beta$-relaxation is generally attributed to spatially local processes involving the rattling motion of a…
VRAM requirements for transformer models scale quadratically with context length due to the self-attention mechanism. In this paper we modify the decoder-only transformer, replacing self-attention with InAttention, which scales linearly…
Understanding whether attention mechanisms converge to the kernel regime is foundational to the validity of influence functions for transformer accountability. Exact NTK characterization of softmax attention is precluded by its exponential…
In long-range percolation on $\mathbb{Z}^d$, points $x$ and $y$ are connected by an edge with probability $1-\exp(-\beta\|x-y\|^{-d-\alpha})$, where $\alpha>0$ is fixed and $\beta \geq 0$ is a parameter. As $d$ and $\alpha$ vary, the model…