IBM 2016 年英语会话电话语音识别系统
计算与语言
2016-06-23 v2
摘要
我们描述了一系列声学与语言建模技术,这些技术将我们的英语会话电话 LVCSR 系统在 Hub5 2000 评测测试集的 Switchboard 子集上的词错误率降低到了创纪录的 6.6%。在声学方面,我们使用了三种强模型的分数融合:具有 maxout 激活的循环网络、具有 3x3 核的极深卷积网络,以及在 FMLLR 和 i-vector 特征上运行的双向长短期记忆网络。在语言建模方面,我们使用了更新后的模型 M 与层次神经网络语言模型(LMs)。
引用
@article{arxiv.1604.08242,
title = {The IBM 2016 English Conversational Telephone Speech Recognition System},
author = {George Saon and Tom Sercu and Steven Rennie and Hong-Kwang J. Kuo},
journal= {arXiv preprint arXiv:1604.08242},
year = {2016}
}
备注
Submitted to Interspeech 2016