中文

用于语音增强的带噪声感知编码器的变分自编码器

音频与语音处理 2021-05-18 v1 机器学习 声音

摘要

近来,一种生成式变分自编码器(VAE)被提出用于语音增强以建模语音统计特性。然而,该方法仅在训练阶段使用干净语音,使得估计对噪声存在尤为敏感,尤其在低信噪比(SNR)下。为提升 VAE 的鲁棒性,我们提出在训练阶段通过使用在含噪-干净语音对上训练的噪声感知编码器来引入噪声信息。我们在不同噪声环境与声学条件下的真实录音上,使用两个不同噪声数据集评估了我们的方法。我们表明,所提出的噪声感知 VAE 在不过增加模型参数数量的情况下,在整体失真方面优于标准 VAE。同时,我们证明我们的模型比有监督前馈深度神经网络(DNN)更能泛化到未见过的噪声条件。此外,我们展示了模型性能对含噪-干净语音训练数据规模缩减的鲁棒性。

关键词

引用

@article{arxiv.2102.08706,
  title  = {Variational Autoencoder for Speech Enhancement with a Noise-Aware Encoder},
  author = {Huajian Fang and Guillaume Carbajal and Stefan Wermter and Timo Gerkmann},
  journal= {arXiv preprint arXiv:2102.08706},
  year   = {2021}
}

备注

ICASSP 2021. (c) 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works