YOLO-Stutter: 端到端的区域感知言语不流畅检测
音频与语音处理
2024-09-17 v3 人工智能
计算与语言
摘要
不流畅言语检测是障碍言语分析和口语学习的瓶颈。当前最先进的模型由基于规则的系统主导,这些系统缺乏效率和鲁棒性,并且对模板设计敏感。在本文中,我们提出了YOLO-Stutter:一种首个以时间精确方式检测不流畅的端到端方法。YOLO-Stutter将不完美的语音-文本对齐作为输入,随后通过空间特征聚合器和时间依赖提取器来执行区域感知的边界和类别预测。我们还引入了两个不流畅语料库,VCTK-Stutter和VCTK-TTS,它们模拟了自然口语不流畅,包括重复、阻塞、缺失、替换和延长。我们的端到端方法在模拟数据和真实失语症语音上,以最少的可训练参数数量实现了最先进的性能。代码和数据集已在 https://github.com/rorizzz/YOLO-Stutter 开源。
引用
@article{arxiv.2408.15297,
title = {YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection},
author = {Xuanru Zhou and Anshul Kashyap and Steve Li and Ayati Sharma and Brittany Morin and David Baquirin and Jet Vonk and Zoe Ezzes and Zachary Miller and Maria Luisa Gorno Tempini and Jiachen Lian and Gopala Krishna Anumanchipalli},
journal= {arXiv preprint arXiv:2408.15297},
year = {2024}
}
备注
Interspeech 2024