非洲语言的大词汇量语音识别:多语言建模与自监督学习
计算与语言
2022-10-05 v2 声音
音频与语音处理
摘要
在非洲使用的 2000 多种语言中,几乎都没有广泛可用的自动语音识别系统,且所需数据也仅对少数语言可用。我们试验了两种可能为非洲语言提供大词汇量语音识别途径的技术:多语言建模与自监督学习。我们收集了可用的开源数据并为 15 种语言采集了数据,并使用这些技术训练了实验模型。我们的结果表明,在多语言端到端模型中汇集可用的少量数据,以及在无监督数据上进行预训练,有助于改善许多非洲语言的语音识别质量。
引用
@article{arxiv.2208.03067,
title = {Large vocabulary speech recognition for languages of Africa: multilingual modeling and self-supervised learning},
author = {Sandy Ritchie and You-Chi Cheng and Mingqing Chen and Rajiv Mathews and Daan van Esch and Bo Li and Khe Chai Sim},
journal= {arXiv preprint arXiv:2208.03067},
year = {2022}
}