English

Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

Audio and Speech Processing 2023-03-15 v2 Artificial Intelligence Sound

Abstract

There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling. However, existing methods on any-speaker adaptive TTS have achieved unsatisfactory performance, due to their suboptimal accuracy in mimicking the target speakers' styles. In this work, we present Grad-StyleSpeech, which is an any-speaker adaptive TTS framework that is based on a diffusion model that can generate highly natural speech with extremely high similarity to target speakers' voice, given a few seconds of reference speech. Grad-StyleSpeech significantly outperforms recent speaker-adaptive TTS baselines on English benchmarks. Audio samples are available at https://nardien.github.io/grad-stylespeech-demo.

Keywords

Cite

@article{arxiv.2211.09383,
  title  = {Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models},
  author = {Minki Kang and Dongchan Min and Sung Ju Hwang},
  journal= {arXiv preprint arXiv:2211.09383},
  year   = {2023}
}

Comments

ICASSP 2023

R2 v1 2026-06-28T06:06:01.971Z