中文

记忆以泛化:高维线性回归中插值必要性研究

机器学习 2022-06-17 v2 机器学习 统计理论 统计理论

摘要

我们考察过参数化模型中插值的必要性,即在机器学习问题中达到最优预测风险需要(近似)插值训练数据。具体地,我们考虑简单的过参数化线性回归 y=Xθ+wy = X \theta + w,其中随机设计 XRn×dX \in \mathbb{R}^{n \times d} 在比例渐近 d/nγ(1,)d/n \to \gamma \in (1, \infty) 下。我们精确刻画了该设定下预测(测试)误差如何必然随训练误差缩放。该刻画的一个推论是:当标签噪声方差 σ20\sigma^2 \to 0 时,任何训练误差至少达到 cσ4\mathsf{c}\sigma^4(对某常数 c\mathsf{c})的估计量必然次优,并将承受至少随训练误差线性增长的额外预测误差。因此,最优性能需要将训练数据拟合到远高于问题固有噪声基底的精度。

关键词

引用

@article{arxiv.2202.09889,
  title  = {Memorize to Generalize: on the Necessity of Interpolation in High Dimensional Linear Regression},
  author = {Chen Cheng and John Duchi and Rohith Kuditipudi},
  journal= {arXiv preprint arXiv:2202.09889},
  year   = {2022}
}

备注

32 pages; accepted to the 35th Annual Conference on Learning Theory (COLT) 2022