English

Hybrid Y-Net Architecture for Singing Voice Separation

Sound 2023-03-07 v1 Machine Learning Audio and Speech Processing

Abstract

This research paper presents a novel deep learning-based neural network architecture, named Y-Net, for achieving music source separation. The proposed architecture performs end-to-end hybrid source separation by extracting features from both spectrogram and waveform domains. Inspired by the U-Net architecture, Y-Net predicts a spectrogram mask to separate vocal sources from a mixture signal. Our results demonstrate the effectiveness of the proposed architecture for music source separation with fewer parameters. Overall, our work presents a promising approach for improving the accuracy and efficiency of music source separation.

Keywords

Cite

@article{arxiv.2303.02599,
  title  = {Hybrid Y-Net Architecture for Singing Voice Separation},
  author = {Rashen Fernando and Pamudu Ranasinghe and Udula Ranasinghe and Janaka Wijayakulasooriya and Pantaleon Perera},
  journal= {arXiv preprint arXiv:2303.02599},
  year   = {2023}
}

Comments

Submitted for EUSIPCO23: 5 Pages, 7 figures

R2 v1 2026-06-28T09:01:50.181Z