English

Controllable and Interpretable Singing Voice Decomposition via Assem-VC

Audio and Speech Processing 2021-10-26 v1 Sound

Abstract

We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize the singing voice of the target speaker. In conclusion, we made a perfectly synced duet with the user's singing voice and the target singer's converted singing voice.

Keywords

Cite

@article{arxiv.2110.12676,
  title  = {Controllable and Interpretable Singing Voice Decomposition via Assem-VC},
  author = {Kang-wook Kim and Junhyeok Lee},
  journal= {arXiv preprint arXiv:2110.12676},
  year   = {2021}
}

Comments

Accepted to NeurIPS Workshop on ML for Creativity and Design 2021 (Oral)

R2 v1 2026-06-24T07:08:59.491Z