English

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models

Machine Learning 2026-05-11 v1 Information Theory math.IT Spectral Theory

Abstract

We study the spectral properties of sample covariance matrices constructed from pooled sequence representations, where token embeddings are drawn from a fixed two-class Gaussian mixture table and pooled via (fixed) attention weights. Working in the high-dimensional regime d,V,Nd,V,N\to\infty with d/Vδd/V\to\delta and d/Nγd/N\to\gamma, we derive exact characterizations of the limiting eigenvalue distribution, outlier eigenvalues, and eigenvector alignment with the hidden signal. The bulk spectrum follows a non-Marchenko--Pastur law given by the free multiplicative convolution κ(MPδMPγ)\kappa(MP_\delta\boxtimes MP_\gamma), reflecting the finite vocabulary structure. Signal recovery undergoes two successive BBP-type phase transitions characterized by the scalars: δ,γ,α=wRw\delta,\gamma,\alpha=w^{\top} R w and κ=w2\kappa=\|w\|^2, where ww denotes the attention pooling weights and RR the positional correlation matrix. An aftermath of our analysis demonstrates that the optimal attention weights maximizing the signal-to-noise ratio α/κ\alpha/\kappa are given by the (normalized) top eigenvector of RR, and we show (as a particular case of our analysis) that parameter-free causal self-attention with τ/d\tau/d score scaling yields deterministic harmonic weights that improve signal recovery over mean pooling whenever early tokens carry more signal. Extensive simulations confirm sharp agreement between theory and finite-dimensional experiments.

Keywords

Cite

@article{arxiv.2605.06826,
  title  = {How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models},
  author = {Mohamed El Amine Seddik},
  journal= {arXiv preprint arXiv:2605.06826},
  year   = {2026}
}
R2 v1 2026-07-01T12:56:02.404Z