English

Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization

Sound 2019-05-01 v3 Audio and Speech Processing Machine Learning

Abstract

In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neural network for modeling the speech spectro-temporal content. The parameters of this supervised model are learned using the framework of variational autoencoders. The noisy recording environment is supposed to be unknown, so the noise spectro-temporal modeling remains unsupervised and is based on non-negative matrix factorization (NMF). We develop a Monte Carlo expectation-maximization algorithm and we experimentally show that the proposed approach outperforms its NMF-based counterpart, where speech is modeled using supervised NMF.

Keywords

Cite

@article{arxiv.1811.06713,
  title  = {Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization},
  author = {Simon Leglaive and Laurent Girin and Radu Horaud},
  journal= {arXiv preprint arXiv:1811.06713},
  year   = {2019}
}

Comments

5 pages, 2 figures, audio examples and code available online at https://team.inria.fr/perception/icassp-2019-mvae/

R2 v1 2026-06-23T05:17:53.773Z