复值循环变分自编码器及其在语音增强中的应用
音频与语音处理
2024-10-28 v2
摘要
作为变分自编码器(VAE)的扩展,复值 VAE 使用复高斯分布对潜变量与数据建模。本文提出一种复值循环 VAE 框架,其中具体采用了复值循环神经网络与 L1 重构损失。首先,为考虑语音信号的时间特性,本文在复值 VAE 框架中引入复值循环神经网络。此外,该框架中使用 L1 损失作为重构损失。为示例说明该复值生成模型在语音处理中的使用,本文选取语音增强作为具体应用。实验基于 TIMIT 数据集。结果表明,所提方法在语音可懂度与信号质量的客观指标上均有提升。
引用
@article{arxiv.2204.02195,
title = {Complex Recurrent Variational Autoencoder with Application to Speech Enhancement},
author = {Yuying Xie and Thomas Arildsen and Zheng-Hua Tan},
journal= {arXiv preprint arXiv:2204.02195},
year = {2024}
}
备注
This work has been submitted to the IEEE for possible publication