中文

面向减少语音训练数据需求的口语理解系统构建方法

计算与语言 2022-03-02 v1 声音 音频与语音处理

摘要

缺乏标注了口语理解(SLU)所需标签的语音数据,通常是构建能够直接处理语音输入的端到端(E2E)系统的主要障碍。相比之下,通常可获得大量带合适标签的文本数据。在本文中,我们提出一种新颖的文本表示与训练方法,使得能够有效利用这些文本资源构建 E2E SLU 系统。借助极少量的额外语音,我们表明这些模型可进一步优化,达到接近基于完整语音数据集构建的同类系统的水平。我们使用三个不同的 SLU 数据集,在意图和实体任务上展示了所提方法的有效性。仅使用文本训练时,所提系统可达到完整语音训练所能实现性能的 90%。仅增加 10% 的语音数据,这些模型便显著提升至完整性能的 97%。

关键词

引用

@article{arxiv.2203.00006,
  title  = {Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems},
  author = {Samuel Thomas and Hong-Kwang J. Kuo and Brian Kingsbury and George Saon},
  journal= {arXiv preprint arXiv:2203.00006},
  year   = {2022}
}

备注

\c{opyright}2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. arXiv admin note: text overlap with arXiv:2202.13155