English

URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement

Audio and Speech Processing 2024-09-25 v1 Sound

Abstract

The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.

Keywords

Cite

@article{arxiv.2406.04660,
  title  = {URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement},
  author = {Wangyou Zhang and Robin Scheibler and Kohei Saijo and Samuele Cornell and Chenda Li and Zhaoheng Ni and Anurag Kumar and Jan Pirklbauer and Marvin Sach and Shinji Watanabe and Tim Fingscheidt and Yanmin Qian},
  journal= {arXiv preprint arXiv:2406.04660},
  year   = {2024}
}

Comments

6 pages, 3 figures, 3 tables. Accepted by Interspeech 2024. An extended version of the accepted manuscript with appendix