中文

Yeah, Un, Oh:基于语音活动投影微调的连续实时背频预测

计算与语言 2025-02-06 v2 人机交互 声音 音频与语音处理

摘要

在人类对话中,诸如“yeah”和“oh”等简短背频言语在促进顺畅且引人入胜的对话中发挥着至关重要的作用。这些背频信号在不中断说话者的前提下传递注意力和理解,因而其准确预测对于创建更自然的对话代理人至关重要。本文提出了一种 novel 方法,用于基于微调语音活动投影(Voice Activity Projection,VAP)的实时连续背频预测。虽然现有方法依赖基于回合或人工平衡的数据集,但我们的 approach 在不平衡的真实世界数据集上,以连续且帧级的方式预测背频的时机和类型。我们首先在通用对话语料库上预训练 VAP 模型以捕捉对话动态,然后在专注于背频行为的 specialized 数据集上进行微调。实验结果表明,我们的模型在时机和类型预测任务中均优于基线方法,在真实环境中实现了稳健的实时性能。这一研究为更具响应性和人类化的对话系统提供了有前景的步骤,对于交互式语音对话应用(如虚拟助手和机器人)具有深远意义。

关键词

引用

@article{arxiv.2410.15929,
  title  = {Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection},
  author = {Koji Inoue and Divesh Lala and Gabriel Skantze and Tatsuya Kawahara},
  journal= {arXiv preprint arXiv:2410.15929},
  year   = {2025}
}

备注

This paper has been accepted for presentation at the main conference of 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025) and represents the author's version of the work