English

AeGAN: Time-Frequency Speech Denoising via Generative Adversarial Networks

Audio and Speech Processing 2020-12-29 v3 Machine Learning Neural and Evolutionary Computing Sound Machine Learning

Abstract

Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in realistic crowded environments. Thus, speech enhancement is a valuable building block in ASR systems and other applications such as hearing aids, smartphones and teleconferencing systems. In this paper, a generative adversarial network (GAN) based framework is investigated for the task of speech enhancement, more specifically speech denoising of audio tracks. A new architecture based on CasNet generator and an additional feature-based loss are incorporated to get realistically denoised speech phonetics. Finally, the proposed framework is shown to outperform other learning and traditional model-based speech enhancement approaches.

Keywords

Cite

@article{arxiv.1910.12620,
  title  = {AeGAN: Time-Frequency Speech Denoising via Generative Adversarial Networks},
  author = {Sherif Abdulatif and Karim Armanious and Karim Guirguis and Jayasankar T. Sajeev and Bin Yang},
  journal= {arXiv preprint arXiv:1910.12620},
  year   = {2020}
}

Comments

5 pages, 4 figures and 2 Tables. Accepted in EUSIPCO 2020

R2 v1 2026-06-23T11:57:03.398Z