一种深度神经网络的新训练框架
机器学习
2021-03-26 v5 计算机视觉与模式识别
摘要
知识蒸馏是将大模型的知识迁移到小模型的过程。在此过程中,小模型学习大模型的泛化能力,并保持接近大模型的性能。知识蒸馏提供了一种迁移模型知识的训练手段,便于模型部署并加速推理。然而,以往的蒸馏方法需要预训练的教师模型,仍带来计算和存储开销。本文提出一种称为自蒸馏(Self Distillation, SD)的新型通用训练框架。我们通过列举其在多种任务和基准数据集上的性能提升来证明方法的有效性。
引用
@article{arxiv.2103.07350,
title = {A New Training Framework for Deep Neural Network},
author = {Zhenyan Hou and Wenxuan Fan},
journal= {arXiv preprint arXiv:2103.07350},
year = {2021}
}
备注
Withdraw this paper for internal review. Because we were not familiar with the use of arXiv, our initial manuscript was uploaded by mistake and we found many inappropriate and unmodified parts of it. I am sorry to say that this work still needs to be further completed and we do not intend to use it for publication