English

Sams-Net: A Sliced Attention-based Neural Network for Music Source Separation

Audio and Speech Processing 2020-05-20 v4 Information Retrieval Machine Learning Sound

Abstract

Convolutional Neural Network (CNN) or Long short-term memory (LSTM) based models with the input of spectrogram or waveforms are commonly used for deep learning based audio source separation. In this paper, we propose a Sliced Attention-based neural network (Sams-Net) in the spectrogram domain for the music source separation task. It enables spectral feature interactions with multi-head attention mechanism, achieves easier parallel computing and has a larger receptive field compared with LSTMs and CNNs respectively. Experimental results on the MUSDB18 dataset show that the proposed method, with fewer parameters, outperforms most of the state-of-the-art DNN-based methods.

Keywords

Cite

@article{arxiv.1909.05746,
  title  = {Sams-Net: A Sliced Attention-based Neural Network for Music Source Separation},
  author = {Tingle Li and Jiawei Chen and Haowen Hou and Ming Li},
  journal= {arXiv preprint arXiv:1909.05746},
  year   = {2020}
}

Comments

Submitted to Interspeech 2020

R2 v1 2026-06-23T11:13:38.846Z