通过增量训练和数据过滤实现多域ASR的更好半监督学习
计算与语言
2025-10-09 v1 声音
音频与语音处理
摘要
针对标签数据稀缺的特定领域ASR模型微调问题。但经常可用 unlabeled audio 和来自相关领域的 labeled 数据。我们提出一种增量式半监督学习管道,首先整合小规模 in-domain 标签数据集和来自 closely related 领域的辅助数据集,实现约4%的相对提升。随后应用基于多模型共识或命名实体识别(NER)的过滤策略用于选择和迭代细化伪标签,显示出相较于随机选择在性能饱和方面更慢。我们在多域 Wow 呼叫中心和 Fisher 英语语料库上进行评估,单步微调的性能优于单一步骤微调。共识-based 过滤法在 Wow 上提供最高可达22.3%的相对提升,在 Fisher 上提供24.8%的相对提升。NER是第二佳筛选器,虽在计算成本较低时仍能提供具竞争力的性能。
引用
@article{arxiv.2506.04981,
title = {Better Semi-supervised Learning for Multi-domain ASR Through Incremental Retraining and Data Filtering},
author = {Andres Carofilis and Pradeep Rangappa and Srikanth Madikeri and Shashi Kumar and Sergio Burdisso and Jeena Prakash and Esau Villatoro-Tello and Petr Motlicek and Bidisha Sharma and Kadri Hacioglu and Shankar Venkatesan and Saurabh Vyas and Andreas Stolcke},
journal= {arXiv preprint arXiv:2506.04981},
year = {2025}
}
备注
Accepted at Interspeech 2025, Netherlands