Related papers: Scaling Limits of Long-Context Transformers
Large language models have an exceptional capability to incorporate new information in a contextual manner. However, the full potential of such an approach is often restrained due to a limitation in the effective context length. One…
We study the scaling properties of critical particle systems confined by a potential. Using renormalization-group arguments, we show that their critical behavior can be cast in the form of a trap-size scaling, resembling finite-size scaling…
We consider two models of one-dimensional discrete random Schrodinger operators (H_n \psi)_l ={\psi}_{l-1}+{\psi}_{l +1}+v_l {\psi}_l, {\psi}_0={\psi}_{n+1}=0 in the cases v_k=\sigma {\omega}_k/\sqrt{n} and v_k=\sigma {\omega}_k/ \sqrt{k}.…
We present an approximate attention mechanism named HyperAttention to address the computational challenges posed by the growing complexity of long contexts used in Large Language Models (LLMs). Recent work suggests that in the worst-case…
Softmax attention defines an interaction through $d_h$ head dimensions, but not all dimensions carry equal weight once real text passes through. We decompose the attention logit field into a learned component and a generated component and…
Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal…
We explore the internal mechanisms of how bias emerges in large language models (LLMs) when provided with ambiguous comparative prompts: inputs that compare or enforce choosing between two or more entities without providing clear context…
We investigate the short time quantum critical dynamics in the imaginary time relaxation processes of finite size systems. Universal scaling behaviors exist in the imaginary time evolution and in particular, the system undergoes a critical…
We present a unifying, consistent, finite-size-scaling picture for percolation theory bringing it into the framework of a general, renormalization-group-based, scaling scheme for systems above their upper critical dimensions $d_c$.…
We report the results of a molecular dynamics simulation of a supercooled binary Lennard-Jones mixture. By plotting the self intermediate scattering functions vs. rescaled time, we find a master curve in the $\beta$-relaxation regime. This…
Large language models (LLMs) are increasingly deployed in privacy-critical and personalization-oriented scenarios, yet the role of context length in shaping privacy leakage and personalization effectiveness remains largely unexplored. We…
In a previous paper we found that in the random field Ising model at zero temperature in three dimensions the correlation length is not self-averaging near the critical point and that the violation of self-averaging is maximal. This is due…
In the Ising model on the simple cubic lattice, we describe the inverse temperature $\beta$ and other quantities relevant for the computation of critical quantities in terms of a dimensionless squared mass $M$. The critical behaviors of…
We consider the problem of estimating inverse temperature parameter $\beta$ of an $n$-dimensional truncated Ising model using a single sample. Given a graph $G = (V,E)$ with $n$ vertices, a truncated Ising model is a probability…
The short-range correlations are considered for a two-dimensional hard-core boson model on square lattice within Bethe approximation for the clusters consisting of two and four sites. Explicit equations are derived for the critical…
We consider an infinite-dimensional stochastic clustering model on $\mathbb{R}$. In discrete time, each point of a unit-intensity simple point process moves halfway toward either of its left or right neighbors, chosen uniformly at random.…
Transformers have proven highly effective across modalities, but standard softmax attention scales quadratically with sequence length, limiting long context modeling. Linear attention mitigates this by approximating attention with kernel…
This thesis investigates two key phenomena in large language models (LLMs): in-context learning (ICL) and model collapse. We study ICL in a linear transformer with tied weights trained on linear regression tasks, and show that minimising…
We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for continuous sequence-to-sequence functions on compact domains. Our…
We study the scaling limit of the rank-one truncation of various beta ensemble generalizations of classical unitary/orthogonal random matrices: the circular beta ensemble, the real orthogonal beta ensemble, and the circular Jacobi beta…