Atss-Net: Target Speaker Separation via Attention-based Neural Network
Audio and Speech Processing
2020-05-20 v1 Sound
Abstract
Recently, Convolutional Neural Network (CNN) and Long short-term memory (LSTM) based models have been introduced to deep learning-based target speaker separation. In this paper, we propose an Attention-based neural network (Atss-Net) in the spectrogram domain for the task. It allows the network to compute the correlation between each feature parallelly, and using shallower layers to extract more features, compared with the CNN-LSTM architecture. Experimental results show that our Atss-Net yields better performance than the VoiceFilter, although it only contains half of the parameters. Furthermore, our proposed model also demonstrates promising performance in speech enhancement.
Cite
@article{arxiv.2005.09200,
title = {Atss-Net: Target Speaker Separation via Attention-based Neural Network},
author = {Tingle Li and Qingjian Lin and Yuanyuan Bao and Ming Li},
journal= {arXiv preprint arXiv:2005.09200},
year = {2020}
}
Comments
Submitted to Interspeech 2020