基于注意力的序列到序列语音识别模型:LibriSpeech 上 SOTA 系统的开发及其在非母语英语中的应用
计算与语言
2018-11-07 v2 机器学习
声音
音频与语音处理
摘要
近期研究表明,基于注意力的序列到序列模型如 Listen, Attend, and Spell (LAS) 在各种任务上取得了与最先进 ASR 系统相当的结果。在本文中,我们描述了此类系统的开发,并展示其在两项任务上的性能:首先,我们在 LibriSpeech 英语数据的 test clean 子集上实现了 3.43% 的新的最先进词错误率;其次,在非母语英语语音(包括朗读语音与自发语音)上,我们相比用最新 Kaldi recipe 构建的传统系统获得了极具竞争力的结果。
引用
@article{arxiv.1810.13088,
title = {Attention-based sequence-to-sequence model for speech recognition: development of state-of-the-art system on LibriSpeech and its application to non-native English},
author = {Yan Yin and Ramon Prieto and Bin Wang and Jianwei Zhou and Yiwei Gu and Yang Liu and Hui Lin},
journal= {arXiv preprint arXiv:1810.13088},
year = {2018}
}