探索 ASR 模型中各层的异质特性以实现更高效训练
机器学习
2022-02-07 v2 计算与语言
声音
音频与语音处理
摘要
基于 Transformer 的架构已成为研究其过参数化及层间非均匀重要性的对象。将这些方法应用于自动语音识别(Automatic Speech Recognition, ASR),我们证明了最先进的 Conformer 模型通常具有多个环境层。我们研究了这些层在不同训练轮次和模型规模下的稳定性,提出可使用组归一化而不破坏其形成,并考察了它们与每层模型权重更新的相关性。最后,我们将这些发现应用于联邦学习(Federated Learning),通过按层重要性实施联邦丢弃(Federated Dropout)来改进训练过程。这使我们能够在不降低质量的前提下减小客户端所优化模型的大小,并展示了未来探索的潜力。
引用
@article{arxiv.2110.04267,
title = {Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient Training},
author = {Lillian Zhou and Dhruv Guliani and Andreas Kabel and Giovanni Motta and Françoise Beaufays},
journal= {arXiv preprint arXiv:2110.04267},
year = {2022}
}
备注
\c{opyright} 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works