Earnings-21:面向真实场景ASR的实用基准
摘要
常用语音语料库不足以充分挑战学术与商业ASR系统。特别是,语音语料库缺乏用于详细分析和词错率(WER)测量所需的元数据。为此,我们提出Earnings-21,一个包含来自九个不同金融板块的实体密集语音、时长39小时的财报电话会议语料库。该语料库旨在对真实场景中的ASR系统进行基准测试,并特别关注命名实体识别。我们对四个商业ASR模型、两个用开源工具构建的内部模型以及一个开源LibriSpeech模型进行基准测试,并讨论它们在Earnings-21上性能的差异。利用我们近期发布的fstalign工具,我们坦诚分析了各模型在不同划分下的识别能力。我们的分析发现,某些命名实体识别(NER)类别的ASR准确率较低,这对转录文本的理解和使用构成显著障碍。Earnings-21衔接了学术与商业ASR系统评估,并为真实世界音频上的实体建模与词错率研究的深入提供了可能。
引用
@article{arxiv.2104.11348,
title = {Earnings-21: A Practical Benchmark for ASR in the Wild},
author = {Miguel Del Rio and Natalie Delworth and Ryan Westerman and Michelle Huang and Nishchal Bhandari and Joseph Palakapilly and Quinten McNamara and Joshua Dong and Piotr Zelasko and Miguel Jette},
journal= {arXiv preprint arXiv:2104.11348},
year = {2022}
}
备注
Accepted to INTERSPEECH 2021. June 15 2021: Addressing the comments of reviewers and updating the results of our internal ESPNet model. The results do not change our conclusions. April 28th, 2021: We found and resolved an issue in our experimental evaluation that scored the LibriSpeech model at ~20% worse relative WER than the actual WER. The updated results do not affect our conclusions