English

Subspace Hybrid MVDR Beamforming for Augmented Hearing

Audio and Speech Processing 2023-12-01 v1 Sound Signal Processing

Abstract

Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their dynamics. However, in the context of augmented reality audio using head-worn microphone arrays, the acoustic scenarios encountered are often far from straightforward. The design of robust, high-performance, adaptive beamformers for such scenarios is an on-going challenge. This is due to the violation of the typically required assumptions on the noise field caused by, for example, rapid variations resulting from complex acoustic environments, and/or rotations of the listener's head. This work proposes a multi-channel speech enhancement algorithm which utilises the adaptability of signal-dependent beamformers while still benefiting from the computational efficiency and robust performance of signal-independent super-directive beamformers. The algorithm has two stages. (i) The first stage is a hybrid beamformer based on a dictionary of weights corresponding to a set of noise field models. (ii) The second stage is a wide-band subspace post-filter to remove any artifacts resulting from (i). The algorithm is evaluated using both real-world recordings and simulations of a cocktail-party scenario. Noise suppression, intelligibility and speech quality results show a significant performance improvement by the proposed algorithm compared to the baseline super-directive beamformer. A data-driven implementation of the noise field dictionary is shown to provide more noise suppression, and similar speech intelligibility and quality, compared to a parametric dictionary.

Keywords

Cite

@article{arxiv.2311.18689,
  title  = {Subspace Hybrid MVDR Beamforming for Augmented Hearing},
  author = {Sina Hafezi and Alastair H. Moore and Pierre H. Guiraud and Patrick A. Naylor and Jacob Donley and Vladimir Tourbabin and Thomas Lunner},
  journal= {arXiv preprint arXiv:2311.18689},
  year   = {2023}
}

Comments

14 pages, 10 figures, submitted for IEEE/ACM Transactions on Audio, Speech, and Language Processing on 23-Nov-2023

R2 v1 2026-06-28T13:37:13.262Z