English

Blind Speech Separation and Dereverberation using Neural Beamforming

Sound 2021-11-08 v2 Machine Learning Audio and Speech Processing

Abstract

In this paper, we present the Blind Speech Separation and Dereverberation (BSSD) network, which performs simultaneous speaker separation, dereverberation and speaker identification in a single neural network. Speaker separation is guided by a set of predefined spatial cues. Dereverberation is performed by using neural beamforming, and speaker identification is aided by embedding vectors and triplet mining. We introduce a frequency-domain model which uses complex-valued neural networks, and a time-domain variant which performs beamforming in latent space. Further, we propose a block-online mode to process longer audio recordings, as they occur in meeting scenarios. We evaluate our system in terms of Scale Independent Signal to Distortion Ratio (SI-SDR), Word Error Rate (WER) and Equal Error Rate (EER).

Keywords

Cite

@article{arxiv.2103.13443,
  title  = {Blind Speech Separation and Dereverberation using Neural Beamforming},
  author = {Lukas Pfeifenberger and Franz Pernkopf},
  journal= {arXiv preprint arXiv:2103.13443},
  year   = {2021}
}

Comments

13 pages, 9 figures

R2 v1 2026-06-24T00:31:54.398Z