English

Karaoker: Alignment-free singing voice synthesis with speech training data

Audio and Speech Processing 2022-09-30 v2 Machine Learning Sound

Abstract

Existing singing voice synthesis models (SVS) are usually trained on singing data and depend on either error-prone time-alignment and duration features or explicit music score information. In this paper, we propose Karaoker, a multispeaker Tacotron-based model conditioned on voice characteristic features that is trained exclusively on spoken data without requiring time-alignments. Karaoker synthesizes singing voice and transfers style following a multi-dimensional template extracted from a source waveform of an unseen singer/speaker. The model is jointly conditioned with a single deep convolutional encoder on continuous data including pitch, intensity, harmonicity, formants, cepstral peak prominence and octaves. We extend the text-to-speech training objective with feature reconstruction, classification and speaker identification tasks that guide the model to an accurate result. In addition to multitasking, we also employ a Wasserstein GAN training scheme as well as new losses on the acoustic model's output to further refine the quality of the model.

Keywords

Cite

@article{arxiv.2204.04127,
  title  = {Karaoker: Alignment-free singing voice synthesis with speech training data},
  author = {Panos Kakoulidis and Nikolaos Ellinas and Georgios Vamvoukakis and Konstantinos Markopoulos and June Sig Sung and Gunu Jho and Pirros Tsiakoulis and Aimilios Chalamandaris},
  journal= {arXiv preprint arXiv:2204.04127},
  year   = {2022}
}

Comments

Accepted to INTERSPEECH 2022

R2 v1 2026-06-24T10:42:34.652Z