English

Neural Vocoders as Speech Enhancers

Sound 2025-01-24 v1 Audio and Speech Processing

Abstract

Speech enhancement (SE) and neural vocoding are traditionally viewed as separate tasks. In this work, we observe them under a common thread: the rank behavior of these processes. This observation prompts two key questions: \textit{Can a model designed for one task's rank degradation be adapted for the other?} and \textit{Is it possible to address both tasks using a unified model?} Our empirical findings demonstrate that existing speech enhancement models can be successfully trained to perform vocoding tasks, and a single model, when jointly trained, can effectively handle both tasks with performance comparable to separately trained models. These results suggest that speech enhancement and neural vocoding can be unified under a broader framework of speech restoration. Code: https://github.com/Andong-Li-speech/Neural-Vocoders-as-Speech-Enhancers.

Keywords

Cite

@article{arxiv.2501.13465,
  title  = {Neural Vocoders as Speech Enhancers},
  author = {Andong Li and Zhihang Sun and Fengyuan Hao and Xiaodong Li and Chengshi Zheng},
  journal= {arXiv preprint arXiv:2501.13465},
  year   = {2025}
}

Comments

6 pages, 3 figures