English

Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis

Sound 2020-11-11 v1 Computation and Language

Abstract

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choice of neural vocoder for waveform synthesis, as well as acoustic configurations used for mel spectrograms and final audio output. We find that fine-tuning a multi-speaker model from found audiobook data that has passed a simple quality threshold can improve naturalness and similarity to unseen target speakers of synthetic speech. Additionally, we find that listeners can discern between a 16kHz and 24kHz sampling rate, and that WaveRNN produces output waveforms of a comparable quality to WaveNet, with a faster inference time.

Keywords

Cite

@article{arxiv.2011.04839,
  title  = {Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis},
  author = {Erica Cooper and Xin Wang and Yi Zhao and Yusuke Yasuda and Junichi Yamagishi},
  journal= {arXiv preprint arXiv:2011.04839},
  year   = {2020}
}

Comments

Technical report