English

Scaling strategies for on-device low-complexity source separation with Conv-Tasnet

Sound 2023-03-07 v1 Machine Learning Audio and Speech Processing

Abstract

Recently, several very effective neural approaches for single-channel speech separation have been presented in the literature. However, due to the size and complexity of these models, their use on low-resource devices, e.g. for hearing aids, and earphones, is still a challenge and established solutions are not available yet. Although approaches based on either pruning or compressing neural models have been proposed, the design of a model architecture suitable for a certain application domain often requires heuristic procedures not easily portable to different low-resource platforms. Given the modular nature of the well-known Conv-Tasnet speech separation architecture, in this paper we consider three parameters that directly control the overall size of the model, namely: the number of residual blocks, the number of repetitions of the separation blocks and the number of channels in the depth-wise convolutions, and experimentally evaluate how they affect the speech separation performance. In particular, experiments carried out on the Libri2Mix show that the number of dilated 1D-Conv blocks is the most critical parameter and that the usage of extra-dilation in the residual blocks allows reducing the performance drop.

Keywords

Cite

@article{arxiv.2303.03005,
  title  = {Scaling strategies for on-device low-complexity source separation with Conv-Tasnet},
  author = {Mohamed Nabih Ali and Francesco Paissan and Daniele Falavigna and Alessio Brutti},
  journal= {arXiv preprint arXiv:2303.03005},
  year   = {2023}
}
R2 v1 2026-06-28T09:03:02.041Z