English

Music Source Separation Based on a Lightweight Deep Learning Framework (DTTNET: DUAL-PATH TFC-TDF UNET)

Audio and Speech Processing 2024-03-20 v2 Sound

Abstract

Music source separation (MSS) aims to extract 'vocals', 'drums', 'bass' and 'other' tracks from a piece of mixed music. While deep learning methods have shown impressive results, there is a trend toward larger models. In our paper, we introduce a novel and lightweight architecture called DTTNet, which is based on Dual-Path Module and Time-Frequency Convolutions Time-Distributed Fully-connected UNet (TFC-TDF UNet). DTTNet achieves 10.12 dB cSDR on 'vocals' compared to 10.01 dB reported for Bandsplit RNN (BSRNN) but with 86.7% fewer parameters. We also assess pattern-specific performance and model generalization for intricate audio patterns.

Keywords

Cite

@article{arxiv.2309.08684,
  title  = {Music Source Separation Based on a Lightweight Deep Learning Framework (DTTNET: DUAL-PATH TFC-TDF UNET)},
  author = {Junyu Chen and Susmitha Vekkot and Pancham Shukla},
  journal= {arXiv preprint arXiv:2309.08684},
  year   = {2024}
}

Comments

Accepted for ICASSP 2024. Additional experiments can be found in the published version on IEEE Xplore

R2 v1 2026-06-28T12:23:02.212Z