中文

某些声音过于常见:利用 Common Voice 数据集构建公平语音识别系统

音频与语音处理 2023-06-07 v1 计算与语言 机器学习 声音

摘要

得益于自监督学习等神经网络训练的新进展,自动语音识别(ASR)系统变得愈发高效。然而,众所周知它们对某些群体不公平,例如带有口音的人。在这项工作中,我们使用法语 Common Voice 数据集来量化预训练的 wav2vec 2.0 模型对若干人口统计群体的偏差。通过在多种固定规模、精心构建的训练集上微调预训练模型,我们展示了说话人多样性的重要性。我们还对 Common Voice 语料库进行了深入分析,并识别出该数据集用户应注意的重要缺陷。

关键词

引用

@article{arxiv.2306.03773,
  title  = {Some voices are too common: Building fair speech recognition systems using the Common Voice dataset},
  author = {Lucas Maison and Yannick Estève},
  journal= {arXiv preprint arXiv:2306.03773},
  year   = {2023}
}

备注

5 pages, 3 figures. Accepted to Interspeech 2023