An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
Abstract
Self-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language adaptation in TTS systems remains an open problem. This paper explores the language adaptation capability of ZMM-TTS, a recent SSL-based multilingual TTS system proposed in our previous work. We conducted experiments on 12 languages using limited data with various fine-tuning configurations. We demonstrate that the similarity in phonetics between the pre-training and target languages, as well as the language category, affects the target language's adaptation performance. Additionally, we find that the fine-tuning dataset size and number of speakers influence adaptability. Surprisingly, we also observed that using paired data for fine-tuning is not always optimal compared to audio-only data. Beyond speech intelligibility, our analysis covers speaker similarity, language identification, and predicted MOS.
Keywords
Cite
@article{arxiv.2406.08911,
title = {An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios},
author = {Cheng Gong and Erica Cooper and Xin Wang and Chunyu Qiang and Mengzhe Geng and Dan Wells and Longbiao Wang and Jianwu Dang and Marc Tessier and Aidan Pine and Korin Richmond and Junichi Yamagishi},
journal= {arXiv preprint arXiv:2406.08911},
year = {2024}
}
Comments
Accepted to Interspeech 2024