过参数化模型中的核机制与丰富机制
机器学习
2020-02-26 v3 机器学习
摘要
近期的一系列工作在“核机制”下研究过参数化神经网络,即网络在训练期间表现得如同核化线性预测器,因此使用梯度下降进行训练等价于寻找最小 RKHS 范数解。这与另外一些研究形成对比,这些研究展示了在过参数化多层网络上使用梯度下降如何能够引入并非 RKHS 范数的丰富隐式偏置。基于 Chizat 和 Bach 的观察,我们展示了初始化的尺度如何控制“核”(又称惰性)机制与“丰富”(又称主动)机制之间的转换,以及它如何影响多层齐次模型的泛化性质。我们针对一个简单的两层模型提供了完整且详细的分析,该模型已经展现出核机制与丰富机制之间有趣且有意义的转换,并且我们针对更复杂的矩阵分解模型和多层非线性网络展示了这种转换。
引用
@article{arxiv.1906.05827,
title = {Kernel and Rich Regimes in Overparametrized Models},
author = {Blake Woodworth and Suriya Gunasekar and Pedro Savarese and Edward Moroshko and Itay Golan and Jason Lee and Daniel Soudry and Nathan Srebro},
journal= {arXiv preprint arXiv:1906.05827},
year = {2020}
}
备注
This paper has been substantially modified, updated, and expanded with additional content (arXiv:2002.09277). To avoid confusion with already existing citations, we are withdrawing the old version of this article