English

Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations

Sound 2024-02-05 v1 Machine Learning Audio and Speech Processing

Abstract

In this paper, we propose a singing voice synthesis model, Karaoker-SSL, that is trained only on text and speech data as a typical multi-speaker acoustic model. It is a low-resource pipeline that does not utilize any singing data end-to-end, since its vocoder is also trained on speech data. Karaoker-SSL is conditioned by self-supervised speech representations in an unsupervised manner. We preprocess these representations by selecting only a subset of their task-correlated dimensions. The conditioning module is indirectly guided to capture style information during training by multi-tasking. This is achieved with a Conformer-based module, which predicts the pitch from the acoustic model's output. Thus, Karaoker-SSL allows singing voice synthesis without reliance on hand-crafted and domain-specific features. There are also no requirements for text alignments or lyrics timestamps. To refine the voice quality, we employ a U-Net discriminator that is conditioned on the target speaker and follows a Diffusion GAN training scheme.

Keywords

Cite

@article{arxiv.2402.01520,
  title  = {Low-Resource Cross-Domain Singing Voice Synthesis via Reduced Self-Supervised Speech Representations},
  author = {Panos Kakoulidis and Nikolaos Ellinas and Georgios Vamvoukakis and Myrsini Christidou and Alexandra Vioni and Georgia Maniati and Junkwang Oh and Gunu Jho and Inchul Hwang and Pirros Tsiakoulis and Aimilios Chalamandaris},
  journal= {arXiv preprint arXiv:2402.01520},
  year   = {2024}
}

Comments

Accepted to IEEE ICASSP SASB 2024

R2 v1 2026-06-28T14:36:01.761Z