中文

改进端到端模型用于口语理解中的集合预测

计算与语言 2022-01-31 v1 机器学习 声音 音频与语音处理

摘要

口语理解(SLU)系统的目标是确定输入语音信号的含义,不同于旨在产生逐字转录的语音识别。端到端(E2E)语音建模的进展使得仅需在语义实体上训练成为可能,而语义实体的收集远比逐字转录廉价。我们关注这一集合预测问题,其中实体顺序未指定。使用两类 E2E 模型——RNN 转导器和基于注意力的编码器-解码器,我们表明当训练实体序列按口语顺序排列时这些模型效果最佳。为改进实体口语顺序未知时的 E2E SLU 模型,我们提出了一种新颖的数据增强技术以及一种基于隐式注意力的对齐方法来推断口语顺序。RNN-T 的 F1 分数显著提升了超过 11%,基于注意力的编码器-解码器 SLU 模型提升了约 2%,优于先前报道的结果。

关键词

引用

@article{arxiv.2201.12105,
  title  = {Improving End-to-End Models for Set Prediction in Spoken Language Understanding},
  author = {Hong-Kwang J. Kuo and Zoltan Tuske and Samuel Thomas and Brian Kingsbury and George Saon},
  journal= {arXiv preprint arXiv:2201.12105},
  year   = {2022}
}

备注

ICASSP \c{opyright}2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works