中文

一种用于意图分类增强的数据增强方法及其在口语对话数据集上的应用

计算与语言 2022-02-22 v1 音频与语音处理

摘要

意图分类器对虚拟智能体系统的成功运行至关重要。在语音激活系统中尤为如此,其中数据可能含有噪声且用户意图存在诸多模糊指向。在运行开始前,这些分类器通常缺乏真实世界训练数据。主动学习是用于辅助标注大量已收集用户输入的常用方法。然而,该方法需要大量人工标注工时。我们提出最近邻得分改进(NNSI)算法用于自动数据选择与标注。NNSI通过自动选择高度模糊样本并以高准确度标注,降低了对人工标注的需求。这是通过整合来自语义相似文本样本组的分类器输出来实现的。标注样本随后可加入训练集以提升分类器准确度。我们在两个大规模真实语音对话系统上演示了NNSI的使用。结果评估表明,我们的方法能够以高准确度选择并标注有用样本。将这些新样本加入训练数据显著改进了分类器,并将错误率降低多达10%。

关键词

引用

@article{arxiv.2202.10137,
  title  = {A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets},
  author = {Zvi Kons and Aharon Satt and Hong-Kwang Kuo and Samuel Thomas and Boaz Carmeli and Ron Hoory and Brian Kingsbury},
  journal= {arXiv preprint arXiv:2202.10137},
  year   = {2022}
}

备注

\c{opyright} 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works