Neural Vocoders as Speech Enhancers
Abstract
Speech enhancement (SE) and neural vocoding are traditionally viewed as separate tasks. In this work, we observe them under a common thread: the rank behavior of these processes. This observation prompts two key questions: \textit{Can a model designed for one task's rank degradation be adapted for the other?} and \textit{Is it possible to address both tasks using a unified model?} Our empirical findings demonstrate that existing speech enhancement models can be successfully trained to perform vocoding tasks, and a single model, when jointly trained, can effectively handle both tasks with performance comparable to separately trained models. These results suggest that speech enhancement and neural vocoding can be unified under a broader framework of speech restoration. Code: https://github.com/Andong-Li-speech/Neural-Vocoders-as-Speech-Enhancers.
Cite
@article{arxiv.2501.13465,
title = {Neural Vocoders as Speech Enhancers},
author = {Andong Li and Zhihang Sun and Fengyuan Hao and Xiaodong Li and Chengshi Zheng},
journal= {arXiv preprint arXiv:2501.13465},
year = {2025}
}
Comments
6 pages, 3 figures