基于语速注意力的时长预测用于语速控制 TTS
音频与语音处理
2023-10-16 v1
摘要
随着高质量语音合成的出现,控制语音各种韵律属性的兴趣大增。语速是建模语音表现力的重要属性。在本工作中,我们提出一种控制非自回归 TTS 语速的新方法。我们通过在时长预测器内部以语速为条件来实现这一点,从而实现隐式语速控制。我们通过以各种语速因子合成音频并测量语速控制合成语音的质量,展示了该方法的优势。此外,我们研究了训练数据的语速分布对有效语速控制的影响。最后,我们对基线预训练 TTS 模型进行微调以获得语速控制 TTS。我们提供各种分析以展示所提方法的优势,以及客观和主观指标。我们发现,相较于基线,所提方法在多个语速因子下具有更高的主观分数和更低的语速误差。
引用
@article{arxiv.2310.08846,
title = {Speaking rate attention-based duration prediction for speed control TTS},
author = {Jesuraj Bandekar and Sathvik Udupa and Abhayjeet Singh and Anjali Jayakumar and Deekshitha G and Sandhya Badiger and Saurabh Kumar and Pooja VH and Prasanta Kumar Ghosh},
journal= {arXiv preprint arXiv:2310.08846},
year = {2023}
}
备注
\c{opyright} 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works