English
Related papers

Related papers: Scaling Limits of Long-Context Transformers

200 papers

Large language models have an exceptional capability to incorporate new information in a contextual manner. However, the full potential of such an approach is often restrained due to a limitation in the effective context length. One…

Computation and Language · Computer Science 2023-12-01 Szymon Tworkowski , Konrad Staniszewski , Mikołaj Pacek , Yuhuai Wu , Henryk Michalewski , Piotr Miłoś

We study the scaling properties of critical particle systems confined by a potential. Using renormalization-group arguments, we show that their critical behavior can be cast in the form of a trap-size scaling, resembling finite-size scaling…

Statistical Mechanics · Physics 2013-05-29 Massimo Campostrini , Ettore Vicari

We consider two models of one-dimensional discrete random Schrodinger operators (H_n \psi)_l ={\psi}_{l-1}+{\psi}_{l +1}+v_l {\psi}_l, {\psi}_0={\psi}_{n+1}=0 in the cases v_k=\sigma {\omega}_k/\sqrt{n} and v_k=\sigma {\omega}_k/ \sqrt{k}.…

Probability · Mathematics 2013-08-02 Evgenij Kritchevski , Benedek Valko , Balint Virag

We present an approximate attention mechanism named HyperAttention to address the computational challenges posed by the growing complexity of long contexts used in Large Language Models (LLMs). Recent work suggests that in the worst-case…

Machine Learning · Computer Science 2023-12-04 Insu Han , Rajesh Jayaram , Amin Karbasi , Vahab Mirrokni , David P. Woodruff , Amir Zandieh

Softmax attention defines an interaction through $d_h$ head dimensions, but not all dimensions carry equal weight once real text passes through. We decompose the attention logit field into a learned component and a generated component and…

Computation and Language · Computer Science 2026-04-09 Wonsuk Lee

Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal…

Machine Learning · Computer Science 2024-11-19 Yuhong Chou , Man Yao , Kexin Wang , Yuqi Pan , Ruijie Zhu , Yiran Zhong , Yu Qiao , Jibin Wu , Bo Xu , Guoqi Li

We explore the internal mechanisms of how bias emerges in large language models (LLMs) when provided with ambiguous comparative prompts: inputs that compare or enforce choosing between two or more entities without providing clear context…

Computation and Language · Computer Science 2024-10-31 Rishabh Adiga , Besmira Nushi , Varun Chandrasekaran

We investigate the short time quantum critical dynamics in the imaginary time relaxation processes of finite size systems. Universal scaling behaviors exist in the imaginary time evolution and in particular, the system undergoes a critical…

Strongly Correlated Electrons · Physics 2017-09-20 Yu-Rong Shu , Shuai Yin , Dao-Xin Yao

We present a unifying, consistent, finite-size-scaling picture for percolation theory bringing it into the framework of a general, renormalization-group-based, scaling scheme for systems above their upper critical dimensions $d_c$.…

Statistical Mechanics · Physics 2017-05-16 Ralph Kenna , Bertrand Berche

We report the results of a molecular dynamics simulation of a supercooled binary Lennard-Jones mixture. By plotting the self intermediate scattering functions vs. rescaled time, we find a master curve in the $\beta$-relaxation regime. This…

Condensed Matter · Physics 2009-10-22 Walter Kob , Hans C. Andersen

Large language models (LLMs) are increasingly deployed in privacy-critical and personalization-oriented scenarios, yet the role of context length in shaping privacy leakage and personalization effectiveness remains largely unexplored. We…

Machine Learning · Computer Science 2026-02-17 Shangding Gu

In a previous paper we found that in the random field Ising model at zero temperature in three dimensions the correlation length is not self-averaging near the critical point and that the violation of self-averaging is maximal. This is due…

Statistical Mechanics · Physics 2007-05-23 Giorgio Parisi , Marco Picco , Nicolas Sourlas

In the Ising model on the simple cubic lattice, we describe the inverse temperature $\beta$ and other quantities relevant for the computation of critical quantities in terms of a dimensionless squared mass $M$. The critical behaviors of…

High Energy Physics - Lattice · Physics 2015-08-25 Hirofumi Yamada

We consider the problem of estimating inverse temperature parameter $\beta$ of an $n$-dimensional truncated Ising model using a single sample. Given a graph $G = (V,E)$ with $n$ vertices, a truncated Ising model is a probability…

Machine Learning · Computer Science 2026-02-17 Rohan Chauhan , Ioannis Panageas

The short-range correlations are considered for a two-dimensional hard-core boson model on square lattice within Bethe approximation for the clusters consisting of two and four sites. Explicit equations are derived for the critical…

Statistical Mechanics · Physics 2022-01-05 E. L. Spevak , A. S. Moskvin , Yu. D. Panov

We consider an infinite-dimensional stochastic clustering model on $\mathbb{R}$. In discrete time, each point of a unit-intensity simple point process moves halfway toward either of its left or right neighbors, chosen uniformly at random.…

Probability · Mathematics 2026-03-10 Partha S. Dey , S. Rasoul Etesami , Aditya S. Gopalan

Transformers have proven highly effective across modalities, but standard softmax attention scales quadratically with sequence length, limiting long context modeling. Linear attention mitigates this by approximating attention with kernel…

Machine Learning · Computer Science 2026-02-10 Ashkan Shahbazi , Chayne Thrash , Yikun Bai , Keaton Hamm , Navid NaderiAlizadeh , Soheil Kolouri

This thesis investigates two key phenomena in large language models (LLMs): in-context learning (ICL) and model collapse. We study ICL in a linear transformer with tied weights trained on linear regression tasks, and show that minimising…

Artificial Intelligence · Computer Science 2026-01-06 Josef Ott

We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on compact domains. Our…

Machine Learning · Computer Science 2025-12-17 Jerry Yao-Chieh Hu , Hude Liu , Hong-Yu Chen , Weimin Wu , Han Liu

We study the scaling limit of the rank-one truncation of various beta ensemble generalizations of classical unitary/orthogonal random matrices: the circular beta ensemble, the real orthogonal beta ensemble, and the circular Jacobi beta…

Probability · Mathematics 2023-10-24 Yun Li , Benedek Valkó