Related papers: Scaling Limits of Long-Context Transformers
Handling long-context sequences efficiently remains a significant challenge in large language models (LLMs). Existing methods for token selection in sequence extrapolation either employ a permanent eviction strategy or select tokens by…
We have extended through beta^{23} the high-temperature expansion of the second field derivative of the susceptibility for Ising models of general spin, with nearest-neighbor interactions, on the simple cubic and the body-centered cubic…
Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting…
We introduce exclusive self attention (XSA), a simple modification of self attention (SA) that improves Transformer's sequence modeling performance. The key idea is to constrain attention to capture only information orthogonal to the…
In Ising model on the simple cubic lattice, we describe the inverse temperature \beta in terms of the bare-mass M and study its critical behavior by the use of delta expansion from high temperature or large M side. In the vicinity of…
We study the out-of-equilibrium behavior of statistical systems along critical relaxational flows arising from instantaneous quenches of the temperature $T$ to the critical point $T_c$, starting from equilibrium conditions at time $t=0$. In…
Identifying words that impact a task's performance more than others is a challenge in natural language processing. Transformers models have recently addressed this issue by incorporating an attention mechanism that assigns greater attention…
We propose the first method to show theoretical limitations for one-layer softmax transformers with arbitrarily many precision bits (even infinite). We establish those limitations for three tasks that require advanced reasoning. The first…
Transformer-based scientific foundation models are increasingly deployed in high-stakes settings, but current architectures give deterministic outputs and provide limited support for calibrated predictive uncertainty. We propose Stochastic…
Efficient attention mechanisms enable long-context transformers but often miss globally important tokens, degrading modeling quality. We introduce a pre-scoring framework that assigns a query-independent global importance prior to keys…
The success of self-attention lies in its ability to capture long-range dependencies and enhance context understanding, but it is limited by its computational complexity and challenges in handling sequential data with inherent…
In-context learning with attention enables large neural networks to make context-specific predictions by selectively focusing on relevant examples. Here, we adapt this idea to supervised learning procedures such as lasso regression and…
We present a new unified theory of critical finite-size scaling for lattice statistical mechanical models with periodic boundary conditions above the upper critical dimension. Our theory is based on recent mathematically rigorous results…
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This…
We conduct athermal simulations of freely-cooling, viscous soft spheres around the jamming transition density \phi_{J}, and find evidence for a growing length \xi(t) that governs relaxation to mechanical equilibrium. \xi(t) is manifest in…
We analyze the scaling parameter, extracted from the fidelity for two different ground states, for the one-dimensional quantum Ising model in a transverse field near the critical point. It is found that, in the thermodynamic limit, the…
Mechanistic interpretability assumes that circuit analysis becomes harder as models scale. We challenge this assumption by showing that the attention architecture matters more than parameter count. Studying three circuit types across Pythia…
Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head…
We present a unified view of finite-size scaling (FSS) in dimension d above the upper critical dimension, for both free and periodic boundary conditions. We find that the modified FSS proposed some time ago to allow for violation of…
We study the distribution of finite size pseudo-critical points in a one-dimensional random quantum magnet with a quantum phase transition described by an infinite randomness fixed point. Pseudo-critical points are defined in three…