English

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Sound 2025-06-23 v1 Audio and Speech Processing

Abstract

To explore the potential advantages of utilizing spatial cues from images for generating stereo singing voices with room reverberation, we introduce VS-Singer, a vision-guided model designed to produce stereo singing voices with room reverberation from scene images. VS-Singer comprises three modules: firstly, a modal interaction network integrates spatial features into text encoding to create a linguistic representation enriched with spatial information. Secondly, the decoder employs a consistency Schr\"odinger bridge to facilitate one-step sample generation. Moreover, we utilize the SFE module to improve the consistency of audio-visual matching. To our knowledge, this study is the first to combine stereo singing voice synthesis with visual acoustic matching within a unified framework. Experimental results demonstrate that VS-Singer can effectively generate stereo singing voices that align with the scene perspective in a single step.

Keywords

Cite

@article{arxiv.2506.16020,
  title  = {VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge},
  author = {Zijing Zhao and Kai Wang and Hao Huang and Ying Hu and Liang He and Jichen Yang},
  journal= {arXiv preprint arXiv:2506.16020},
  year   = {2025}
}

Comments

Accepted by Interspeech 2025