English
Related papers

Related papers: Scaling Limits of Long-Context Transformers

200 papers

The softmax function is crucial in Transformer attention, which normalizes each row of the attention scores with summation to one, achieving superior performances over other alternative functions. However, the softmax function can face a…

Computation and Language · Computer Science 2025-02-26 Chuanyang Zheng , Yihang Gao , Guoxuan Chen , Han Shi , Jing Xiong , Xiaozhe Ren , Chao Huang , Xin Jiang , Zhenguo Li , Yu Li

The ability to process long contexts is crucial for many natural language processing tasks, yet it remains a significant challenge. While substantial progress has been made in enhancing the efficiency of attention mechanisms, there is still…

Computation and Language · Computer Science 2025-03-06 Konstantin Donhauser , Charles Arnal , Mohammad Pezeshki , Vivien Cabannes , David Lopez-Paz , Kartik Ahuja

There are two independent critical exponents that describe the behavior of systems near their critical point. However, at the critical point only the exponent $\eta$, which describes the decay of the correlation function, is usually…

Statistical Mechanics · Physics 2015-06-25 S. Davatolhagh

The jamming transition of particles with finite-range interactions is characterized by a variety of critical phenomena, including power law distributions of marginal contacts. We numerically study a recently proposed simple model of…

Statistical Mechanics · Physics 2016-01-20 Yoav Kallus

We examine a phase transition in a model of random spatial permutations which originates in a study of the interacting Bose gas. Permutations are weighted according to point positions; the low-temperature onset of the appearance of…

Statistical Mechanics · Physics 2015-05-14 John Kerl

Transformers have become the go-to architecture for language and vision tasks, yet their theoretical properties, especially memorization capacity, remain elusive. This paper investigates the memorization abilities of multi-head attention…

Machine Learning · Computer Science 2024-03-05 Sadegh Mahdavi , Renjie Liao , Christos Thrampoulidis

We simulated site dilute Ising models in $d=3$ dimensions for several lattice sizes $L$. For each $L$ singular thermodynamic quantities $X$ were measured at criticality and their distributions $P(X)$ were determined, for ensembles of…

Disordered Systems and Neural Networks · Physics 2007-05-23 S. Wiseman , E. Domany

Softmax Self-Attention (SSA) is a key component of Transformer architectures. However, when utilised within skipless architectures, which aim to improve representation learning, recent work has highlighted the inherent instability of SSA…

Machine Learning · Computer Science 2026-02-06 Leo Zhang , James Martens

Let $\Lambda=\{\Lambda_0,\Lambda_1,\Lambda_2,\ldots\}$ be the point process that describes the edge scaling limit of either (i) "regular" beta-ensembles with inverse temperature $\beta>0$, or (ii) the top eigenvalues of Wishart or Gaussian…

Probability · Mathematics 2025-12-10 Pierre Yves Gaudreau Lamarre

We identify intrinsic limitations of Rotary Positional Embeddings (RoPE) in Transformer-based long-context language models. Our theoretical analysis abstracts away from the specific content of the context and depends only on its length. We…

Computation and Language · Computer Science 2026-05-18 Yufeng Du , Phillip Harris , Minyang Tian , Eliu A Huerta , Srikanth Ronanki , Subendhu Rongali , Aram Galstyan , Hao Peng

Modern sequence modeling is dominated by two families: Transformers, whose self-attention can access arbitrary elements of the visible sequence, and structured state-space models, which propagate information through an explicit recurrent…

Machine Learning · Computer Science 2026-04-22 Liubomyr Horbatko

The Ising model in two dimensions with the special boundary conditions of Brascamp and Kunz is analysed. Leading and sub-dominant scaling behaviour of the Fisher zeroes are determined exactly. The finite-size scaling, with corrections, of…

Statistical Mechanics · Physics 2009-11-07 W. Janke , R. Kenna

Extreme events can come either from point processes, when the size or energy of the events is above a certain threshold, or from time series, when the intensity of a signal surpasses a threshold value. We are particularly concerned by the…

Statistical Mechanics · Physics 2017-07-26 Alvaro Corral

We analyze the block averaging transformation applied to lattice gas models with short range interaction in the uniqueness region below the critical temperature. We prove weak Gibbsianity of the renormalized measure and convergence of the…

Statistical Mechanics · Physics 2015-05-30 L. Bertini , Emilio N. M. Cirillo , E. Olivieri

Attention mechanism is a central component of the transformer architecture which led to the phenomenal success of large language models. However, the theoretical principles underlying the attention mechanism are poorly understood,…

Machine Learning · Computer Science 2023-12-11 Davoud Ataee Tarzanagh , Yingcong Li , Xuechen Zhang , Samet Oymak

Extensive efforts have been made to boost the performance in the domain of language models by introducing various attention-based transformers. However, the inclusion of linear layers with large dimensions contributes to significant…

Machine Learning · Computer Science 2024-11-19 Priyansh Bhatnagar , Linfeng Wen , Mingu Kang

We study the dynamical response of a system to a sudden change of the tuning parameter $\lambda$ starting (or ending) at the quantum critical point. In particular we analyze the scaling of the excitation probability, number of excited…

Statistical Mechanics · Physics 2010-01-20 C. De Grandi , V. Gritsev , A. Polkovnikov

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks. However, the quadratic complexity of softmax attention remains a central bottleneck that limits their scalability. Alman and Song (NeurIPS…

Machine Learning · Computer Science 2026-03-20 Maryam Aliakbarpour , Vladimir Braverman , Junze Yin , Haochen Zhang

A $d$--dimensional quantum model in the spherical approximation confined to a general geometry of the form $L^{d-d^{\prime}} \times\infty^{d^{\prime}}\times L_{\tau}^{z}$ ($L$--linear space size and $L_{\tau}$--temporal size) and subjected…

Condensed Matter · Physics 2008-02-03 H. Chamati , D. M. Danchev , E. S. Pisanova , N. S. Tonchev

Linear attention reduces the quadratic cost of softmax attention to $\mathcal{O}(T)$, but its memory state grows as $\mathcal{O}(T)$ in Frobenius norm, causing progressive interference between stored associations. We introduce…

Machine Learning · Computer Science 2026-05-13 Vishal Pandey , Gopal Singh