English

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

Sound 2026-05-04 v2

Abstract

Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.

Keywords

Cite

@article{arxiv.2603.02641,
  title  = {Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement},
  author = {Szu-Wei Fu and Rong Chao and Xuesong Yang and Sung-Feng Huang and Ryandhimas E. Zezario and Rauf Nasretdinov and Ante Jukić and Yu Tsao and Yu-Chiang Frank Wang},
  journal= {arXiv preprint arXiv:2603.02641},
  year   = {2026}
}
R2 v1 2026-07-01T11:00:30.168Z