English

Analyzing the impact of speaker localization errors on speech separation for automatic speech recognition

Audio and Speech Processing 2019-10-25 v1

Abstract

We investigate the effect of speaker localization on the performance of speech recognition systems in a multispeaker, multichannel environment. Given the speaker location information, speech separation is performed in three stages. In the first stage, a simple delay-and-sum (DS) beamformer is used to enhance the signal impinging from the speaker location which is then used to estimate a time-frequency mask corresponding to the localized speaker using a neural network. This mask is used to compute the second order statistics and to derive an adaptive beamformer in the third stage. We generated a multichannel, multispeaker, reverberated, noisy dataset inspired from the well studied WSJ0-2mix and study the performance of the proposed pipeline in terms of the word error rate (WER). An average WER of 29.429.4% was achieved using the ground truth localization information and 42.442.4% using the localization information estimated via GCC-PHAT. The signal-to-interference ratio (SIR) between the speakers has a higher impact on the ASR performance, to the extent of reducing the WER by 5959% relative for a SIR increase of 1515 dB. By contrast, increasing the spatial distance to 5050^\circ or more improves the WER by 2323% relative only

Keywords

Cite

@article{arxiv.1910.11114,
  title  = {Analyzing the impact of speaker localization errors on speech separation for automatic speech recognition},
  author = {Sunit Sivasankaran and Emmaneul Vincent and Dominique Fohr},
  journal= {arXiv preprint arXiv:1910.11114},
  year   = {2019}
}

Comments

Submitted to ICASSP 2020

R2 v1 2026-06-23T11:53:42.894Z