某些声音过于常见:利用 Common Voice 数据集构建公平语音识别系统
音频与语音处理
2023-06-07 v1 计算与语言
机器学习
声音
摘要
得益于自监督学习等神经网络训练的新进展,自动语音识别(ASR)系统变得愈发高效。然而,众所周知它们对某些群体不公平,例如带有口音的人。在这项工作中,我们使用法语 Common Voice 数据集来量化预训练的 wav2vec 2.0 模型对若干人口统计群体的偏差。通过在多种固定规模、精心构建的训练集上微调预训练模型,我们展示了说话人多样性的重要性。我们还对 Common Voice 语料库进行了深入分析,并识别出该数据集用户应注意的重要缺陷。
引用
@article{arxiv.2306.03773,
title = {Some voices are too common: Building fair speech recognition systems using the Common Voice dataset},
author = {Lucas Maison and Yannick Estève},
journal= {arXiv preprint arXiv:2306.03773},
year = {2023}
}
备注
5 pages, 3 figures. Accepted to Interspeech 2023