The People's Speech:一个面向商业用途的大规模多样化英语语音识别数据集
机器学习
2021-11-19 v1 机器学习
摘要
The People's Speech 是一个可免费下载的、规模达 30,000 小时且不断增长的监督对话式英语语音识别数据集,依据 CC-BY-SA(含 CC-BY 子集)授权用于学术与商业用途。数据通过搜索互联网上带有现有转写文本且授权适当的音频数据来收集。我们描述了数据收集方法,并在 Apache 2.0 许可下发布了我们的数据收集系统。我们展示了在该数据集上训练的模型在 Librispeech 的 test-clean 测试集上取得了 9.98% 的词错误率。最后,我们讨论了围绕创建一个大规模机器学习语料的合法与伦理问题,以及由 MLCommons 赞助下对该项目持续维护的计划。
引用
@article{arxiv.2111.09344,
title = {The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage},
author = {Daniel Galvez and Greg Diamos and Juan Ciro and Juan Felipe Cerón and Keith Achorn and Anjali Gopi and David Kanter and Maximilian Lam and Mark Mazumder and Vijay Janapa Reddi},
journal= {arXiv preprint arXiv:2111.09344},
year = {2021}
}
备注
Part of 2021 Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks