人类与机器在英语广播新闻语音识别上的表现
摘要
随着深度学习的近期进展,人们相当关注在诸如会话电话语音(CTS)识别等任务上实现接近人类表现的自动语音识别性能。在本文中,我们评估了这些被提出的技术在一个类似且具有挑战性的任务——广播新闻(BN)上的有用性。我们还进行了一系列识别测量,以理解所达成的自动语音识别结果在该任务上距离人类表现有多近。在两个公开可用的 BN 测试集 DEV04F 和 RT04 上,我们使用基于 LSTM 和残差网络的声学模型并结合 n-gram 与神经网络语言模型的语音识别系统,分别取得了 6.5% 和 5.9% 的词错误率。通过在这些测试集上达成新的性能里程碑,我们的实验表明,在 CTS 等其他相关任务上开发的技术可以迁移以取得类似表现。相比之下,在这些测试集上测得的最佳人类识别表现要低得多,分别为 3.6% 和 2.8%,表明该领域仍有空间开发新技术并改进,以达人类表现水平。
引用
@article{arxiv.1904.13258,
title = {English Broadcast News Speech Recognition by Humans and Machines},
author = {Samuel Thomas and Masayuki Suzuki and Yinghui Huang and Gakuto Kurata and Zoltan Tuske and George Saon and Brian Kingsbury and Michael Picheny and Tom Dibert and Alice Kaiser-Schatzlein and Bern Samko},
journal= {arXiv preprint arXiv:1904.13258},
year = {2019}
}
备注
\copyright 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works