eCat:一种用于多说话人TTS与多对多细粒度韵律迁移的端到端模型
音频与语音处理
2023-06-21 v1 声音
摘要
我们提出eCat,一种新颖的端到端多说话人模型,能够:a)生成具有表现力且符合上下文的韵律的长上下文语音;b)在任何一对已见说话人之间执行细粒度韵律迁移。eCat采用两阶段训练方法训练。在第一阶段,模型从语音中以端到端方式学习说话人无关的词语级韵律表示。在第二阶段,我们学习利用文本中可用的上下文信息预测韵律表示。我们将eCat与CopyCat2(一种能够进行细粒度韵律迁移(FPT)和多说话人TTS的模型)进行比较。我们表明,在2种语言、3个地区和7位说话人上,eCat使CopyCat2与人类录音在自然度上的差距平均缩小46.7%,且在FPT中具有更好的目标说话人相似度。我们还将eCat与VITS比较,并展示出统计显著的偏好。
引用
@article{arxiv.2306.11327,
title = {eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer},
author = {Ammar Abbas and Sri Karlapati and Bastian Schnell and Penny Karanasou and Marcel Granero Moya and Amith Nagaraj and Ayman Boustati and Nicole Peinelt and Alexis Moinet and Thomas Drugman},
journal= {arXiv preprint arXiv:2306.11327},
year = {2023}
}
备注
Accepted to be published in the Proceedings of InterSpeech 2023