English

SLOGD: Speaker LOcation Guided Deflation approach to speech separation

Audio and Speech Processing 2019-10-25 v1

Abstract

Speech separation is the process of separating multiple speakers from an audio recording. In this work we propose to separate the sources using a Speaker LOcalization Guided Deflation (SLOGD) approach wherein we estimate the sources iteratively. In each iteration we first estimate the location of the speaker and use it to estimate a mask corresponding to the localized speaker. The estimated source is removed from the mixture before estimating the location and mask of the next source. Experiments are conducted on a reverberated, noisy multichannel version of the well-studied WSJ-2MIX dataset using word error rate (WER) as a metric. The proposed method achieves a WER of 44.244.2%, a 3434% relative improvement over the system without separation and 1717% relative improvement over Conv-TasNet.

Keywords

Cite

@article{arxiv.1910.11131,
  title  = {SLOGD: Speaker LOcation Guided Deflation approach to speech separation},
  author = {Sunit Sivasankaran and Emmanuel Vincent and Dominique Fohr},
  journal= {arXiv preprint arXiv:1910.11131},
  year   = {2019}
}

Comments

Submitted to ICASSP 2020

R2 v1 2026-06-23T11:53:44.778Z