English

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Audio and Speech Processing 2024-06-05 v1 Sound

Abstract

We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-TTS serves as a foundation model for speech generation and excels in speech in-context learning, achieving performance in speaker similarity and naturalness that matches ground truth human speech in both objective and subjective evaluations. With fine-tuning, we achieve even higher subjective scores across these metrics. Seed-TTS offers superior controllability over various speech attributes such as emotion and is capable of generating highly expressive and diverse speech for speakers in the wild. Furthermore, we propose a self-distillation method for speech factorization, as well as a reinforcement learning approach to enhance model robustness, speaker similarity, and controllability. We additionally present a non-autoregressive (NAR) variant of the Seed-TTS model, named Seed-TTSDiT\text{Seed-TTS}_\text{DiT}, which utilizes a fully diffusion-based architecture. Unlike previous NAR-based TTS systems, Seed-TTSDiT\text{Seed-TTS}_\text{DiT} does not depend on pre-estimated phoneme durations and performs speech generation through end-to-end processing. We demonstrate that this variant achieves comparable performance to the language model-based variant and showcase its effectiveness in speech editing. We encourage readers to listen to demos at \url{https://bytedancespeech.github.io/seedtts_tech_report}.

Keywords

Cite

@article{arxiv.2406.02430,
  title  = {Seed-TTS: A Family of High-Quality Versatile Speech Generation Models},
  author = {Philip Anastassiou and Jiawei Chen and Jitong Chen and Yuanzhe Chen and Zhuo Chen and Ziyi Chen and Jian Cong and Lelai Deng and Chuang Ding and Lu Gao and Mingqing Gong and Peisong Huang and Qingqing Huang and Zhiying Huang and Yuanyuan Huo and Dongya Jia and Chumin Li and Feiya Li and Hui Li and Jiaxin Li and Xiaoyang Li and Xingxing Li and Lin Liu and Shouda Liu and Sichao Liu and Xudong Liu and Yuchen Liu and Zhengxi Liu and Lu Lu and Junjie Pan and Xin Wang and Yuping Wang and Yuxuan Wang and Zhen Wei and Jian Wu and Chao Yao and Yifeng Yang and Yuanhao Yi and Junteng Zhang and Qidi Zhang and Shuo Zhang and Wenjie Zhang and Yang Zhang and Zilin Zhao and Dejian Zhong and Xiaobin Zhuang},
  journal= {arXiv preprint arXiv:2406.02430},
  year   = {2024}
}