English

Compressible Softmax-Attended Language under Incompressible Attention

Computation and Language 2026-04-09 v2 Artificial Intelligence

Abstract

Softmax attention defines an interaction through dhd_h head dimensions, but not all dimensions carry equal weight once real text passes through. We decompose the attention logit field into a learned component and a generated component and measure their spectra separately. For all 5,888 KV heads in five transformer language models (124M--7B parameters, four architecture families), the logit energy field E~\tilde{E} reaches 90\% of its variance in 2--11 singular components. The learned interaction matrix WQTWKW_Q^\mathrm{T} W_K needs 38--75 components for the same threshold out of dh64,128d_h \in {64, 128}. The spectral gap is 5--25×\times in effective rank. The compressibility of softmax-attended language is a property of the data, not the frame that analyzes it.

Cite

@article{arxiv.2604.04384,
  title  = {Compressible Softmax-Attended Language under Incompressible Attention},
  author = {Wonsuk Lee},
  journal= {arXiv preprint arXiv:2604.04384},
  year   = {2026}
}

Comments

6 pages

R2 v1 2026-07-01T11:54:53.500Z