通过最小范数解的二层ReLU神经网络先验泛化误差
机器学习
2020-05-08 v3 机器学习
摘要
我们致力于估计由均方误差训练的二层ReLU神经网络(NNs)的\emph{先验}泛化误差,其仅依赖于初始参数与目标函数,研究思路如下。我们首先估计带最小范数解约束的有限宽二层ReLU NN的\emph{先验}泛化误差,\cite{zhang2019type}证明其为线性化(关于参数)有限宽二层NN的等价解。当宽度趋于无穷时,该线性化NN收敛至Neural Tangent Kernel(NTK)机制中的NN \citep{jacot2018neural}。因而,我们可推导NTK机制中二层ReLU NN的\emph{先验}泛化误差。NTK机制中NN与梯度训练的有限宽NN间的距离由\cite{arora2019exact}估计。基于\cite{arora2019exact}的结果,我们的工作证明了二层ReLU NNs的一个\emph{先验}泛化误差界。该估计利用最小范数解固有的隐式偏置,无需损失函数中的额外正则性。该\emph{先验}估计还表明NN不受维数灾难影响,且可在不要求指数级大量神经元的情况下实现小泛化误差。此外,本文提出的研究思路亦可用于研究有限宽网络的其他性质,如后验泛化误差。
引用
@article{arxiv.1912.03011,
title = {A priori generalization error for two-layer ReLU neural network through minimum norm solution},
author = {Zhi-Qin John Xu and Jiwei Zhang and Yaoyu Zhang and Chengchao Zhao},
journal= {arXiv preprint arXiv:1912.03011},
year = {2020}
}
备注
There is a error in this paper that the scale of initialization in this paper is different from the NTK regime. So, the generalization error of the neural network in the NTK regime is baseless