中文

SynParaSpeech:面向语音生成与理解的鸡尾酒数据集的自动合成框架

音频与语音处理 2025-09-30 v3 计算与语言

摘要

鸡尾酒声(paralinguistic sounds),如笑声和叹气,对于合成更真实、更具沉浸感的语音,以及改进语音理解具有关键作用。然而,现有方法通常依赖专有数据集,而公开可获取的资源常常存在语音不完整、时间戳不准确或缺乏现实相关性的问题。为此,我们提出一种自动化框架,用于生成大规模鸡尾酒数据,并据此构建了SynParaSpeech数据集。该数据集包含6类鸡尾酒,总时长达118.75小时,所有数据均来自自然对话语音,并附有精确时间戳。我们的贡献在于首次提出构建大规模鸡尾酒数据集的自动化方法,发布SynParaSpeech语料库,推动语音生成通过更自然的鸡尾酒合成,增强语音理解通过改进鸡尾酒事件检测。该数据集及音频样本已发布于https://github.com/ShawnPi233/SynParaSpeech。

关键词

引用

@article{arxiv.2509.14946,
  title  = {SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding},
  author = {Bingsong Bai and Qihang Lu and Wenbing Yang and Zihan Sun and Yueran Hou and Peilei Jia and Songbai Pu and Ruibo Fu and Yingming Gao and Ya Li and Jun Gao},
  journal= {arXiv preprint arXiv:2509.14946},
  year   = {2025}
}

备注

Submitted to ICASSP 2026. Copyright 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works