English

On the Interpolation Error of Nonlinear Attention versus Linear Regression

Machine Learning 2026-02-27 v2 Machine Learning Statistics Theory Statistics Theory

Abstract

Attention has become the core building block of modern machine learning (ML) by efficiently capturing the long-range dependencies among input tokens. Its inherently parallelizable structure allows for efficient performance scaling with the rapidly increasing size of both data and model parameters. Despite its central role, the theoretical understanding of Attention, especially in the nonlinear setting, is progressing at a more modest pace. This paper provides a precise characterization of the interpolation error for a nonlinear Attention, in the high-dimensional regime where the number of input tokens nn and the embedding dimension pp are both large and comparable. Under a signal-plus-noise data model and for fixed Attention weights, we derive explicit (limiting) expressions for the mean-squared interpolation error. Leveraging recent advances in random matrix theory, we show that nonlinear Attention generally incurs a larger interpolation error than linear regression on random inputs. However, this gap vanishes, and can even be reversed, when the input contains a structured signal, particularly if the Attention weights align with the signal direction. Our theoretical insights are supported by numerical experiments.

Keywords

Cite

@article{arxiv.2506.18656,
  title  = {On the Interpolation Error of Nonlinear Attention versus Linear Regression},
  author = {Zhenyu Liao and Jiaqing Liu and TianQi Hou and Difan Zou and Zenan Ling},
  journal= {arXiv preprint arXiv:2506.18656},
  year   = {2026}
}

Comments

37 pages, 7 figures

R2 v1 2026-07-01T03:29:29.886Z