English

ExARN: self-attending RNN for target speaker extraction

Audio and Speech Processing 2023-03-14 v2 Signal Processing

Abstract

Target speaker extraction is to extract the target speaker, specified by enrollment utterance, in an environment with other competing speakers. Therefore, the task needs to solve two problems, speaker identification and separation, at the same time. In this paper, we combine self-attention and Recurrent Neural Networks (RNN). Further, we exploit various ways to combining different auxiliary information with mixed representations. Experimental results show that our proposed model achieves excellent performance on the task of target speaker extraction.

Keywords

Cite

@article{arxiv.2212.01106,
  title  = {ExARN: self-attending RNN for target speaker extraction},
  author = {Pengjie Shen and Shulin He and Xueliang Zhang},
  journal= {arXiv preprint arXiv:2212.01106},
  year   = {2023}
}

Comments

The overall quality of the article is not good enough

R2 v1 2026-06-28T07:20:20.991Z