中文

迈向语音识别公平性度量:Casual Conversations 数据集转写

音频与语音处理 2021-11-22 v1 声音

摘要

众所周知,许多机器学习系统对特定群体表现出偏见。该问题在人脸识别领域已被广泛研究,但在自动语音识别(ASR)中则少得多。本文给出在“Casual Conversations”上的初步语音识别结果——这是一个公开发布的 846 小时语料库,旨在帮助研究者评估其计算机视觉与音频模型在年龄、性别和肤色等多样元数据上的准确性。整个语料库已人工转写,允许跨这些元数据进行详细的 ASR 评估。我们评估了多个 ASR 模型,包括基于 LibriSpeech 训练的模型、14000 小时转写数据训练的模型,以及超过 200 万小时未转写社交媒体视频训练的模型。所有模型均在某些时刻观测到跨性别与肤色的词错误率显著差异。我们正发布 Casual Conversations 数据集的人工转写文本,以鼓励社区开发多种技术来减少这些统计偏见。

关键词

引用

@article{arxiv.2111.09983,
  title  = {Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions},
  author = {Chunxi Liu and Michael Picheny and Leda Sarı and Pooja Chitkara and Alex Xiao and Xiaohui Zhang and Mark Chou and Andres Alvarado and Caner Hazirbas and Yatharth Saraf},
  journal= {arXiv preprint arXiv:2111.09983},
  year   = {2021}
}

备注

Submitted to ICASSP 2022. Our dataset will be publicly available at (https://ai.facebook.com/datasets/casual-conversations-downloads) for general use. We also would like to note that considering the limitations of our dataset, we limit the use of it for only evaluation purposes (see license agreement)