English

A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance

Machine Learning 2025-12-09 v1 Artificial Intelligence

Abstract

We develop a unified mathematical framework for certified Top-kk attention truncation that quantifies approximation error at both the distribution and output levels. For a single attention distribution PP and its Top-kk truncation P^\hat P, we show that the total-variation distance coincides with the discarded softmax tail mass and satisfies TV(P,P^)=1eKL(P^P)\mathrm{TV}(P,\hat P)=1-e^{-\mathrm{KL}(\hat P\Vert P)}, yielding sharp Top-kk-specific bounds in place of generic inequalities. From this we derive non-asymptotic deterministic bounds -- from a single boundary gap through multi-gap and blockwise variants -- that control TV(P,P^)\mathrm{TV}(P,\hat P) using only the ordered logits. Using an exact head-tail decomposition, we prove that the output error factorizes as Attn(q,K,V)Attnk(q,K,V)2=τμtailμhead2\|\mathrm{Attn}(q,K,V)-\mathrm{Attn}_k(q,K,V)\|_2=\tau\|\mu_{\mathrm{tail}}-\mu_{\mathrm{head}}\|_2 with τ=TV(P,P^)\tau=\mathrm{TV}(P,\hat P), yielding a new head-tail diameter bound Attn(q,K,V)Attnk(q,K,V)2τdiamH,T\|\mathrm{Attn}(q,K,V)-\mathrm{Attn}_k(q,K,V)\|_2\le\tau\,\mathrm{diam}_{H,T} and refinements linking the error to VarP(V)\mathrm{Var}_P(V). Under an i.i.d. Gaussian score model siN(μ,σ2)s_i\sim\mathcal N(\mu,\sigma^2) we derive closed-form tail masses and an asymptotic rule for the minimal kεk_\varepsilon ensuring TV(P,P^)ε\mathrm{TV}(P,\hat P)\le\varepsilon, namely kε/nΦc(σ+Φ1(ε))k_\varepsilon/n\approx\Phi_c(\sigma+\Phi^{-1}(\varepsilon)). Experiments on bert-base-uncased and synthetic logits confirm the predicted scaling of kε/nk_\varepsilon/n and show that certified Top-kk can reduce scored keys by 2-4×\times on average while meeting the prescribed total-variation budget.

Keywords

Cite

@article{arxiv.2512.07647,
  title  = {A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance},
  author = {Georgios Tzachristas and Lei Deng and Ioannis Tzachristas and Gong Zhang and Renhai Chen},
  journal= {arXiv preprint arXiv:2512.07647},
  year   = {2025}
}