English

PITS: Variational Pitch Inference without Fundamental Frequency for End-to-End Pitch-controllable TTS

Audio and Speech Processing 2023-06-07 v3 Machine Learning Sound

Abstract

Previous pitch-controllable text-to-speech (TTS) models rely on directly modeling fundamental frequency, leading to low variance in synthesized speech. To address this issue, we propose PITS, an end-to-end pitch-controllable TTS model that utilizes variational inference to model pitch. Based on VITS, PITS incorporates the Yingram encoder, the Yingram decoder, and adversarial training of pitch-shifted synthesis to achieve pitch-controllability. Experiments demonstrate that PITS generates high-quality speech that is indistinguishable from ground truth speech and has high pitch-controllability without quality degradation. Code, audio samples, and demo are available at https://github.com/anonymous-pits/pits.

Keywords

Cite

@article{arxiv.2302.12391,
  title  = {PITS: Variational Pitch Inference without Fundamental Frequency for End-to-End Pitch-controllable TTS},
  author = {Junhyeok Lee and Wonbin Jung and Hyunjae Cho and Jaeyeon Kim and Jaehwan Kim},
  journal= {arXiv preprint arXiv:2302.12391},
  year   = {2023}
}

Comments

6 pages, preprint