English

Binaural Speech Enhancement Using STOI-Optimal Masks

Audio and Speech Processing 2022-10-03 v1 Signal Processing

Abstract

STOI-optimal masking has been previously proposed and developed for single-channel speech enhancement. In this paper, we consider the extension to the task of binaural speech enhancement in which spatial information is known to be important to speech understanding and therefore should be preserved by the enhancement processing. Masks are estimated for each of the binaural channels individually and a `better-ear listening' mask is computed by choosing the maximum of the two masks. The estimated mask is used to supply probability information about the speech presence in each time-frequency bin to an Optimally-modified Log Spectral Amplitude (OM-LSA) enhancer. We show that using the proposed method for binaural signals with a directional noise not only improves the SNR of the noisy signal but also preserves the binaural cues and intelligibility.

Keywords

Cite

@article{arxiv.2209.15472,
  title  = {Binaural Speech Enhancement Using STOI-Optimal Masks},
  author = {Vikas Tokala and Mike Brookes and Patrick A. Naylor},
  journal= {arXiv preprint arXiv:2209.15472},
  year   = {2022}
}

Comments

Accepted at IWAENC 2022

R2 v1 2026-06-28T02:27:35.889Z