基于潜在特征条件的 GAN 语音增强:低信噪比场景
音频与语音处理
2024-10-18 v1 声音
信号处理
摘要
在不利的信噪比 (SNR) 条件下增强语音质量,仍是判别式深度神经网络 (DNN) 基于方法的一个重要挑战。本文提出 DisCoGAN,即基于判别式模型在低 SNR 场景下预训练的语音增强潜在特征条件的时频域生成对抗网络 (GAN)。我们的方法在性能上优于最新的判别式方法,也超越端到端 (E2E) 训练的 GAN 模型。我们还研究了 various configurations for conditioning the proposed GAN model with the discriminative model 对增强语音质量的影响。
引用
@article{arxiv.2410.13599,
title = {GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning},
author = {Shrishti Saha Shetu and Emanuël A. P. Habets and Andreas Brendel},
journal= {arXiv preprint arXiv:2410.13599},
year = {2024}
}
备注
5 pages, 2 figures