OTF:基于最优传输的监督与自监督学习模型融合用于自动语音识别
音频与语音处理
2023-06-06 v1
摘要
自监督学习(Self-Supervised Learning, SSL)自动语音识别(Automatic Speech Recognition, ASR)模型在低资源场景下相较监督学习(Supervised Learning, SL)模型展现出巨大潜力。然而,在许多工业应用中,当标注数据量增加时,SSL 的优势逐渐减弱。为了在拥有充足标注数据时进一步提升 ASR 性能,我们首先通过分析 SL 与 SSL ASR 模型在识别准确率和优化特性上的互补性,探索了二者结合的潜力。随后,我们提出了一种新颖的基于最优传输的融合(Optimal Transport based Fusion, OTF)方法,用于 SL 与 SSL 模型,且在推理时不引入额外计算开销。具体而言,采用最优传输对逐层权重进行软对齐,将两个不同网络统一为单一网络。在公开 1k 小时英文 LibriSpeech 数据集和内部 2.6k 小时中文数据集上的实验结果表明,OTF 以更低的错误率大幅优于单个模型。
引用
@article{arxiv.2306.02541,
title = {OTF: Optimal Transport based Fusion of Supervised and Self-Supervised Learning Models for Automatic Speech Recognition},
author = {Li Fu and Siqi Li and Qingtao Li and Fangzhu Li and Liping Deng and Lu Fan and Meng Chen and Youzheng Wu and Xiaodong He},
journal= {arXiv preprint arXiv:2306.02541},
year = {2023}
}
备注
Accepted by Interspeech 2023