PRoDeliberation:面向端到端 spoken language understanding 的并行鲁棒 deliberation
计算与语言
2024-06-13 v1 声音
音频与语音处理
摘要
口语理解(SLU)是语音助理的关键组件,由将语音转换为语义解析以执行任务构成。以往的工作探索了使用 deliberation 方法来提高 SLU 模型的质量和鲁棒性,但这些模型仍是自回归的,导致延迟较高。本文引入了 PRoDeliberation,一种新方法,利用基于 Connectionist Temporal Classification 的解码策略以及去噪目标来训练鲁棒的非自回归 deliberation 模型。我们展示,PRoDeliberation 实现了并行解码的延迟降低(相对于自回归模型提高 2-10 倍),同时保留了纠正自回归 deliberation 系统自动语音识别(ASR)误译的能力。我们进一步表明,去噪训练设计使 PRoDeliberation 能够克服小型 ASR 设备的局限性,并对系统各组件的必要性进行了分析。
引用
@article{arxiv.2406.07823,
title = {PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding},
author = {Trang Le and Daniel Lazar and Suyoun Kim and Shan Jiang and Duc Le and Adithya Sagar and Aleksandr Livshits and Ahmed Aly and Akshat Shrivastava},
journal= {arXiv preprint arXiv:2406.07823},
year = {2024}
}