中文

StoryTTS:具备丰富文本表达注释的高度可表达文本语音合成数据集

声音 2024-04-24 v1 计算与语言 音频与语音处理

摘要

虽然语音可表达性在可表达文本语音合成(ETTS)领域长期受到关注,但文本本身固有的可表达性却未得到充分关注,尤其是对于艺术作品的ETTS。本文介绍了StoryTTS,一个包含语音和文本两个维度丰富可表达性的高质量ETTS数据集,源自一档中文叙事节目录音。我们提出了系统且全面的文本可表达性标注框架。通过 linguistics、修辞等方法,我们分析并定义了StoryTTS中与语音相关的文本可表达性,包括五个不同维度。随后我们运用大语言模型,并以少量手动标注示例为提示,进行批量标注。最终得到的语料库包含61小时连续且高度抒情的语音,配有准确的文本转录和丰富的文本可表达性注释。因此,StoryTTS可为未来ETTS研究提供支持,使其充分挖掘丰富的内在文本和语音特征。我们 conducted experiments to validate that TTS models can generate speech with improved expressiveness when integrating with the annotated textual labels in StoryTTS.

关键词

引用

@article{arxiv.2404.14946,
  title  = {StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations},
  author = {Sen Liu and Yiwei Guo and Xie Chen and Kai Yu},
  journal= {arXiv preprint arXiv:2404.14946},
  year   = {2024}
}

备注

Accepted by ICASSP 2024