中文

DiDiSpeech:一个大规模普通话语音语料库

音频与语音处理 2021-02-09 v4

摘要

本文介绍一个新的开源普通话语音语料库,称为DiDiSpeech。它包含来自6000名说话人、采样率为48kHz、约800小时语音数据及相应文本。语料库中所有语音数据均在安静环境中录制,适用于多种语音处理任务,如语音转换、多说话人文本到语音以及自动语音识别。我们使用多种语音任务进行实验并评估性能,表明该语料库用于学术研究与实际应用中均有良好前景。语料库可从 https://outreach.didichuxing.com/research/opendata/ 获取。

关键词

引用

@article{arxiv.2010.09275,
  title  = {DiDiSpeech: A Large Scale Mandarin Speech Corpus},
  author = {Tingwei Guo and Cheng Wen and Dongwei Jiang and Ne Luo and Ruixiong Zhang and Shuaijiang Zhao and Wubo Li and Cheng Gong and Wei Zou and Kun Han and Xiangang Li},
  journal= {arXiv preprint arXiv:2010.09275},
  year   = {2021}
}

备注

5 pages, 2 figures, 11 tables