English

Using IPA-Based Tacotron for Data Efficient Cross-Lingual Speaker Adaptation and Pronunciation Enhancement

Sound 2022-04-01 v2 Machine Learning Audio and Speech Processing

Abstract

Recent neural Text-to-Speech (TTS) models have been shown to perform very well when enough data is available. However, fine-tuning them for new speakers or languages is not straightforward in a low-resource setup. In this paper, we show that by applying minor modifications to a Tacotron model, one can transfer an existing TTS model for new speakers from the same or a different language using only 20 minutes of data. For this purpose, we first introduce a base multi-lingual Tacotron with language-agnostic input, then demonstrate how transfer learning is done for different scenarios of speaker adaptation without exploiting any pre-trained speaker encoder or code-switching technique. We evaluate the transferred model in both subjective and objective ways.

Keywords

Cite

@article{arxiv.2011.06392,
  title  = {Using IPA-Based Tacotron for Data Efficient Cross-Lingual Speaker Adaptation and Pronunciation Enhancement},
  author = {Hamed Hemati and Damian Borth},
  journal= {arXiv preprint arXiv:2011.06392},
  year   = {2022}
}

Comments

Preprint

R2 v1 2026-06-23T20:08:08.597Z