English

BiSinger: Bilingual Singing Voice Synthesis

Audio and Speech Processing 2024-01-10 v3 Machine Learning Sound

Abstract

Although Singing Voice Synthesis (SVS) has made great strides with Text-to-Speech (TTS) techniques, multilingual singing voice modeling remains relatively unexplored. This paper presents BiSinger, a bilingual pop SVS system for English and Chinese Mandarin. Current systems require separate models per language and cannot accurately represent both Chinese and English, hindering code-switch SVS. To address this gap, we design a shared representation between Chinese and English singing voices, achieved by using the CMU dictionary with mapping rules. We fuse monolingual singing datasets with open-source singing voice conversion techniques to generate bilingual singing voices while also exploring the potential use of bilingual speech data. Experiments affirm that our language-independent representation and incorporation of related datasets enable a single model with enhanced performance in English and code-switch SVS while maintaining Chinese song performance. Audio samples are available at https://bisinger-svs.github.io.

Keywords

Cite

@article{arxiv.2309.14089,
  title  = {BiSinger: Bilingual Singing Voice Synthesis},
  author = {Huali Zhou and Yueqian Lin and Yao Shi and Peng Sun and Ming Li},
  journal= {arXiv preprint arXiv:2309.14089},
  year   = {2024}
}

Comments

Accepted by ASRU2023

R2 v1 2026-06-28T12:31:31.757Z