中文

Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models

机器学习 2026-05-15 v1 机器学习 统计理论 统计理论

摘要

We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction. Specifically, we consider the problem of recovering a multi-index polynomial f(x)=h(Ux)f^*(x)=h(Ux), with URr×dU\in\mathbb{R}^{r\times d} and rdr\ll d, from finitely many data/label pairs. Importantly, the target function depends on input xx only through the projection onto an unknown rr-dimensional central subspace. The algorithm we analyze is appealingly simple: fit kernel ridge regression (KRR) to the data and compute the Average Gradient Outer Product (AGOP) from the fitted predictor. Our main results show that under reasonable assumptions the top rr-dimensional eigenspace of AGOP provably recovers the central subspace, even in regimes when the prediction error remains large. Specifically, if the target function ff^* has degree pp^*, it is known that ndpn\asymp d^{p^*} samples are necessary for KRR to achieve accurate prediction. In contrast, we show that if a low degree pp component of ff^* already carries all relevant directions for prediction, subspace recovery occurs in the much lower sample regime ndp+δn\asymp d^{p+\delta} for any δ(0,1)\delta\in(0,1). Our results thus demonstrate a separation between prediction and representation, and provide an explanation for why iterative kernel methods such as Recursive Feature Machines (RFM) can be sample-efficient in practice.

关键词

引用

@article{arxiv.2605.15082,
  title  = {Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models},
  author = {Libin Zhu and Damek Davis and Dmitriy Drusvyatskiy and Maryam Fazel},
  journal= {arXiv preprint arXiv:2605.15082},
  year   = {2026}
}

备注

95 pages, 12 figures