Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
Machine Learning
2026-02-26 v2 Machine Learning
Probability
Statistics Theory
Statistics Theory
Abstract
We study the convergence rate of learning pairwise interactions in single-layer attention-style models, where tokens interact through a weight matrix and a nonlinear activation function. We prove that the minimax rate is , where is the sample size and is the H\"older smoothness of the activation function. Importantly, this rate is independent of the embedding dimension , the number of tokens , and the rank of the weight matrix, provided that . These results highlight a fundamental statistical efficiency of attention-style models, even when the weight matrix and activation are not separately identifiable, and provide a theoretical understanding of attention mechanisms and guidance on training.
Cite
@article{arxiv.2510.11789,
title = {Minimax Rates for Learning Pairwise Interactions in Attention-Style Models},
author = {Shai Zucker and Xiong Wang and Fei Lu and Inbar Seroussi},
journal= {arXiv preprint arXiv:2510.11789},
year = {2026}
}