English

ERNIE-SAT: Speech and Text Joint Pretraining for Cross-Lingual Multi-Speaker Text-to-Speech

Audio and Speech Processing 2022-12-06 v2 Computation and Language Sound

Abstract

Speech representation learning has improved both speech understanding and speech synthesis tasks for single language. However, its ability in cross-lingual scenarios has not been explored. In this paper, we extend the pretraining method for cross-lingual multi-speaker speech synthesis tasks, including cross-lingual multi-speaker voice cloning and cross-lingual multi-speaker speech editing. We propose a speech-text joint pretraining framework, where we randomly mask the spectrogram and the phonemes given a speech example and its transcription. By learning to reconstruct the masked parts of the input in different languages, our model shows great improvements over speaker-embedding-based multi-speaker TTS methods. Moreover, our framework is end-to-end for both the training and the inference without any finetuning effort. In cross-lingual multi-speaker voice cloning and cross-lingual multi-speaker speech editing tasks, our experiments show that our model outperforms speaker-embedding-based multi-speaker TTS methods.

Keywords

Cite

@article{arxiv.2211.03545,
  title  = {ERNIE-SAT: Speech and Text Joint Pretraining for Cross-Lingual Multi-Speaker Text-to-Speech},
  author = {Xiaoran Fan and Chao Pang and Tian Yuan and He Bai and Renjie Zheng and Pengfei Zhu and Shuohuan Wang and Junkun Chen and Zeyu Chen and Liang Huang and Yu Sun and Hua Wu},
  journal= {arXiv preprint arXiv:2211.03545},
  year   = {2022}
}