English

MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset

Audio and Speech Processing 2023-01-03 v1 Artificial Intelligence Computation and Language

Abstract

Text-to-Speech (TTS) synthesis for low-resource languages is an attractive research issue in academia and industry nowadays. Mongolian is the official language of the Inner Mongolia Autonomous Region and a representative low-resource language spoken by over 10 million people worldwide. However, there is a relative lack of open-source datasets for Mongolian TTS. Therefore, we make public an open-source multi-speaker Mongolian TTS dataset, named MnTTS2, for the benefit of related researchers. In this work, we prepare the transcription from various topics and invite three professional Mongolian announcers to form a three-speaker TTS dataset, in which each announcer records 10 hours of speeches in Mongolian, resulting 30 hours in total. Furthermore, we build the baseline system based on the state-of-the-art FastSpeech2 model and HiFi-GAN vocoder. The experimental results suggest that the constructed MnTTS2 dataset is sufficient to build robust multi-speaker TTS models for real-world applications. The MnTTS2 dataset, training recipe, and pretrained models are released at: \url{https://github.com/ssmlkl/MnTTS2}

Keywords

Cite

@article{arxiv.2301.00657,
  title  = {MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset},
  author = {Kailin Liang and Bin Liu and Yifan Hu and Rui Liu and Feilong Bao and Guanglai Gao},
  journal= {arXiv preprint arXiv:2301.00657},
  year   = {2023}
}

Comments

Accepted by NCMMSC'2022 (https://ncmmsc2022.ustc.edu.cn/main.htm)

R2 v1 2026-06-28T07:59:33.236Z