English

DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise

Audio and Speech Processing 2026-05-12 v1

Abstract

Acoustic echo and background noise pose challenges on speech enhancement in hands-free systems and speakerphones. Discriminatively trained end-to-end methods represent a powerful solution for joint acoustic echo control (AEC) and denoising. However, with the advent of generative methods, diffusion-based approaches have seen remarkable performance in speech enhancement tasks. In this work, to the best of our knowledge, we provide the first (still non-causal) diffusion-based AEC model (DiffVQE) that is reproducible in terms of topology, training data, and training framework. So far, without employing diffusion, Microsoft's discriminative DeepVQE model has been shown to excel any of the ICASSP 2023 AEC Challenge entries achieving remarkable performance. Using data from the Interspeech 2025 URGENT Challenge for a diverse, high-quality training dataset, our DiffVQE excels DeepVQE both in echo and noise control performance, as well as in computational complexity and model size.

Keywords

Cite

@article{arxiv.2605.08189,
  title  = {DiffVQE: Hybrid Diffusion Voice Quality Enhancement Under Acoustic Echo and Noise},
  author = {Haljan Lugo Girao and Ernst Seidel and Pejman Mowlaee and Ziyue Zhao and Tim Fingscheidt},
  journal= {arXiv preprint arXiv:2605.08189},
  year   = {2026}
}

Comments

6 pages, 4 figures, submitted to Interspeech 2026

R2 v1 2026-07-01T12:58:30.448Z