大型非线性模型的线性性:切核何时及为何为常数
机器学习
2021-02-23 v3 机器学习
摘要
本文旨在阐明某些神经网络在其宽度趋于无穷时转变为线性的显著现象。我们表明,模型向线性化的转变,等价地,切核(神经切核,NTK)的常数性,源于网络 Hessian 矩阵范数作为网络宽度函数的缩放性质。我们提出了一个通过 Hessian 缩放理解切核常数性的通用框架,适用于标准类的神经网络。我们的分析为常切核现象提供了与广泛接受的“惰性训练”不同的新视角。此外,我们表明向线性化的转变并非宽神经网络的普遍性质,当网络最后一层为非线性时不成立,且对于梯度下降的成功优化也并非必要。
引用
@article{arxiv.2010.01092,
title = {On the linearity of large non-linear models: when and why the tangent kernel is constant},
author = {Chaoyue Liu and Libin Zhu and Mikhail Belkin},
journal= {arXiv preprint arXiv:2010.01092},
year = {2021}
}
备注
accepted as Spotlight in NeurIPS 2020; made correction to proof