English

Universal Speech Enhancement with Regression and Generative Mamba

Sound 2025-10-01 v2 Audio and Speech Processing

Abstract

The Interspeech 2025 URGENT Challenge aimed to advance universal, robust, and generalizable speech enhancement by unifying speech enhancement tasks across a wide variety of conditions, including seven different distortion types and five languages. We present Universal Speech Enhancement Mamba (USEMamba), a state-space speech enhancement model designed to handle long-range sequence modeling, time-frequency structured processing, and sampling frequency-independent feature extraction. Our approach primarily relies on regression-based modeling, which performs well across most distortions. However, for packet loss and bandwidth extension, where missing content must be inferred, a generative variant of the proposed USEMamba proves more effective. Despite being trained on only a subset of the full training data, USEMamba achieved 2nd place in Track 1 during the blind test phase, demonstrating strong generalization across diverse conditions.

Keywords

Cite

@article{arxiv.2505.21198,
  title  = {Universal Speech Enhancement with Regression and Generative Mamba},
  author = {Rong Chao and Rauf Nasretdinov and Yu-Chiang Frank Wang and Ante Jukić and Szu-Wei Fu and Yu Tsao},
  journal= {arXiv preprint arXiv:2505.21198},
  year   = {2025}
}

Comments

Accepted to Interspeech 2025

R2 v1 2026-07-01T02:43:01.468Z