深度学习中优化的几何与隐式正则化
机器学习
2017-05-10 v1
摘要
我们认为优化通过隐式正则化在深度学习模型的泛化中起着关键作用。我们通过证明泛化能力并非由网络规模控制,而是由其他某种隐式控制所主导来做到这一点。我们随后展示了改变经验优化过程如何能够改善泛化,即使实际优化质量未受影响。我们通过研究深度网络参数空间的几何结构,并设计一种适应此几何结构的优化算法来实现这一点。
引用
@article{arxiv.1705.03071,
title = {Geometry of Optimization and Implicit Regularization in Deep Learning},
author = {Behnam Neyshabur and Ryota Tomioka and Ruslan Salakhutdinov and Nathan Srebro},
journal= {arXiv preprint arXiv:1705.03071},
year = {2017}
}
备注
This survey chapter was done as a part of Intel Collaborative Research institute for Computational Intelligence (ICRI-CI) "Why & When Deep Learning works -- looking inside Deep Learning" compendium with the generous support of ICRI-CI. arXiv admin note: substantial text overlap with arXiv:1506.02617