GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
Abstract
We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.
Cite
@article{arxiv.2604.01832,
title = {GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement},
author = {Xiaobin Rong and Yushi Wang and Zheng Wang and Jing Lu},
journal= {arXiv preprint arXiv:2604.01832},
year = {2026}
}
Comments
Awarded 1st place in the URGENT 2026 Challenge (objective phase), accepted by ICASSP 2026