English

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

Audio and Speech Processing 2021-06-17 v1 Computation and Language Sound

Abstract

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs speaker embeddings to identify and extract multiple targets without label permutation ambiguity. To efficiently inform the speaker information to the extraction model, we propose a new speaker conditioning mechanism by designing an additional speaker branch for receiving external speaker embeddings. Experiments on 2-channel WHAMR! data show that the proposed system improves by 9% relative the source separation performance over a strong multi-channel baseline, and it increases the speech recognition accuracy by more than 16% relative over the same baseline.

Keywords

Cite

@article{arxiv.2102.03762,
  title  = {Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism},
  author = {Jisi Zhang and Catalin Zorila and Rama Doddipatla and Jon Barker},
  journal= {arXiv preprint arXiv:2102.03762},
  year   = {2021}
}

Comments

Accepted for ICASSP 2021

R2 v1 2026-06-23T22:54:40.366Z