Zambezi Voice:面向赞比亚语言的多语言语音语料库
计算与语言
2023-06-16 v2 声音
音频与语音处理
摘要
本工作介绍 Zambezi Voice,一个面向赞比亚语言的开源多语言语音资源。它包含两类数据集:无标注的广播新闻与谈话节目音频录音(160 小时),以及由公开文学书籍文本转录的朗读语音组成的标注数据(超 80 小时)。该数据集为语音识别而创建,但可扩展至有监督与无监督学习方法的多语言语音处理研究。据我们所知,这是首个为赞比亚语言创建的多语言语音数据集。我们利用预训练与跨语言迁移学习,通过微调 Wav2Vec2.0 大规模多语言预训练模型来构建端到端(E2E)语音识别基线模型。数据集以知识共享 BY-NC-ND 4.0 许可公开发布,可通过 https://github.com/unza-speech-lab/zambezi-voice 获取。
引用
@article{arxiv.2306.04428,
title = {Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages},
author = {Claytone Sikasote and Kalinda Siaminwe and Stanly Mwape and Bangiwe Zulu and Mofya Phiri and Martin Phiri and David Zulu and Mayumbo Nyirenda and Antonios Anastasopoulos},
journal= {arXiv preprint arXiv:2306.04428},
year = {2023}
}
备注
Accepted at INTERSPEECH 2023. This pre-print version differs slightly from the version accepted to INTERSPEECH 2023: Figure 1 is not included in INTERSPEECH 2023!