中文

集成文本输入用于RNN转导器ASR模型的训练与自适应

计算与语言 2022-03-01 v1 声音 音频与语音处理

摘要

与使用模块化架构、各组件可独立适配新领域的混合自动语音识别(ASR)系统相比,近来的端到端(E2E)ASR系统由于其全神经整体式结构而更难定制。在本文中,我们提出一种用于 E2E ASR 模型的新型文本表示与训练框架。通过该方法,我们展示了训练好的 RNN 转导器(RNN-T)模型的内部 LM 组件可使用纯文本数据得到有效适配。使用语音与文本输入共同训练的 RNN-T 模型,在 NIST Hub5 2000 评测的 Switchboard 和 CallHome 测试集上,相较仅用语音训练的基线模型取得了接近 13% 的词错误率(WER)降低。所提方法的实用性通过将该通用 RNN-T 模型定制到三个独立数据集得到进一步证明。在这些设定下,利用仅来自新领域的非配对文本数据,我们通过这种新颖的 LM 风格定制技术观察到 20-45% 的相对词错误率(WER)降低。

关键词

引用

@article{arxiv.2202.13155,
  title  = {Integrating Text Inputs For Training and Adapting RNN Transducer ASR Models},
  author = {Samuel Thomas and Brian Kingsbury and George Saon and Hong-Kwang J. Kuo},
  journal= {arXiv preprint arXiv:2202.13155},
  year   = {2022}
}

备注

\c{opyright}2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works