Common Voice 中的 Quechua 语音数据集:Puno Quechua 的案例研究
计算与语言
2025-10-17 v1
摘要
资源不足的语言,如契丁语言(Quechua),面临数据和资源的稀缺,阻碍了其在语音技术领域的发展。为此,Common Voice 提供了一个关键机会,促进开放且由社区驱动的语音数据集创建。本文考察了将契丁语言整合到 Common Voice 中的情况。我们详细介绍了当前17种契丁语言, presenting Puno Quechua(ISO 639-3: qxp)作为一个聚焦案例研究,包括语言 onboarding 和来自阅读语音数据和自发语音数据的语料收集。我们的结果表明,Common Voice 现已拥有191.1小时的契丁语音音频(86%已验证),其中 Puno Quechua 贡献了12小时(77%已验证),凸显了 Common Voice 的潜力。我们进一步提出了一项研究议程,解决技术挑战,并为社区参与和原住民数据主权伦理考虑提供指导。我们的工作推动了包容性语音技术和资源不足语言社区的数字赋能。
引用
@article{arxiv.2510.13871,
title = {Quechua Speech Datasets in Common Voice: The Case of Puno Quechua},
author = {Elwin Huaman and Wendi Huaman and Jorge Luis Huaman and Ninfa Quispe},
journal= {arXiv preprint arXiv:2510.13871},
year = {2025}
}
备注
to be published in the 9th Annual International Conference on Information Management and Big Data (SIMBig 2025)