Related papers: Scaling Limits of Long-Context Transformers
Softmax and related normalized response functions are widely used in choice theory, machine learning, and cognitive science. In non-Boolean event structures with overlapping contexts, however, local normalization does not automatically…
Transformers are extremely successful machine learning models whose mathematical properties remain poorly understood. Here, we rigorously characterize the behavior of transformers with hardmax self-attention and normalization sublayers as…
The Transformer, with its scaled dot-product attention mechanism, has become a foundational architecture in modern AI. However, this mechanism is computationally intensive and incurs substantial energy costs. We propose a new Transformer…
Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed attention. The multiple heads learn diverse types of word…
Using the machinery of smooth scaling and coarse-graining of observables, developed recently in the context of so-called fluctuation operators (originally developed by Verbeure et al), we extend this approach to a rigorous renormalisation…
Large language models (LLMs) have demonstrated strong performance on a variety of natural language processing (NLP) tasks. However, they often struggle with long-text sequences due to the ``lost in the middle'' phenomenon. This issue has…
The influence of a thermodynamic constraint on the critical finite-size scaling behavior of three-dimensional Ising and XY models is analyzed by Monte-Carlo simulations. Within the Ising universality class constraints lead to Fisher…
In long-range percolation on $\mathbb{Z}^d$, points $x$ and $y$ are connected by an edge with probability $1-\exp(-\beta\|x-y\|^{-d-\alpha})$, where $\alpha>0$ is fixed and $\beta \geq 0$ is a parameter. As $d$ and $\alpha$ vary, the model…
Efficient long-context modeling remains a critical challenge for natural language processing (NLP), as the time complexity of the predominant Transformer architecture scales quadratically with the sequence length. While state-space models…
Training causal transformers at extreme sequence lengths is bottlenecked by the quadratic time and memory of scaled dot-product attention (SDPA). In this work, we propose Lighthouse Attention, a training-only symmetrical selection-based…
Transformers have emerged as the architecture of choice for many state-of-the-art AI models, showcasing exceptional performance across a wide range of AI applications. However, the memory demands imposed by Transformers limit their ability…
We develop a scaling theory for the finite-size critical behavior of the microcanonical entropy (density of states) of a system with a critically-divergent heat capacity. The link between the microcanonical entropy and the canonical energy…
The distributions $P(X)$ of singular thermodynamic quantities in an ensemble of quenched random samples of linear size $l$ at the critical point $T_c$ are studied by Monte Carlo in two models. Our results confirm predictions of Aharony and…
We investigate the integration of human-like working memory constraints into the Transformer architecture and implement several cognitively inspired attention variants, including fixed-width windows based and temporal decay based attention…
Local-global attention models have recently emerged as compelling alternatives to standard Transformers, promising improvements in both training and inference efficiency. However, the crucial choice of window size presents a Pareto…
Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To help rectify these…
Finite-size scaling is a key tool in statistical physics, used to infer critical behavior in finite systems. Here we use the analogous concept of finite-time scaling to describe the bifurcation diagram at finite times in discrete dynamical…
The finite-size scaling theory for continuous phase transition plays an important role in determining critical point and critical exponents from the size-dependent behaviors of quantities in the thermodynamic limit. For percolation phase…
We have investigated the anomalous scaling behaviour of the Ising model on small-world networks based on 2- and 3-dimensional lattices using Monte Carlo simulations. Our main result is that even at low $p$, the shift in the critical…
We try to design a simple model exhibiting self-organized criticality, which is amenable to a rigorous mathematical analysis. To this end, we modify the generalized Ising Curie-Weiss model by implementing an automatic control of the inverse…