English

Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction

Sound 2023-06-16 v1 Audio and Speech Processing

Abstract

This paper describes a real-time General Speech Reconstruction (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This novel proposed system is a two-stage architecture, in which the speech restoration is performed, and then cascaded by speech enhancement. We propose a complex spectral mapping-based generative adversarial network (CSM-GAN) as the speech restoration module for the first time. For noise suppression and dereverberation, the enhancement module is performed with fullband-wideband parallel processing. On the blind test set of ICASSP 2023 SSI Challenge, the proposed Gesper system, which satisfies the real-time condition, achieves 3.27 P.804 overall mean opinion score (MOS) and 3.35 P.835 overall MOS, ranked 1st in both track 1 and track 2.

Keywords

Cite

@article{arxiv.2306.08454,
  title  = {Gesper: A Restoration-Enhancement Framework for General Speech Reconstruction},
  author = {Wenzhe Liu and Yupeng Shi and Jun Chen and Wei Rao and Shulin He and Andong Li and Yannan Wang and Zhiyong Wu},
  journal= {arXiv preprint arXiv:2306.08454},
  year   = {2023}
}

Comments

Accepted by InterSpeech 2023

R2 v1 2026-06-28T11:04:56.851Z