English

SRTNet: Time Domain Speech Enhancement Via Stochastic Refinement

Sound 2022-11-01 v1 Audio and Speech Processing

Abstract

Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose SRTNet, a novel method for speech enhancement via Stochastic Refinement in complete Time domain. Specifically, we design a joint network consisting of a deterministic module and a stochastic module, which makes up the ``enhance-and-refine'' paradigm. We theoretically demonstrate the feasibility of our method and experimentally prove that our method achieves faster training, faster sampling and higher quality. Our code and enhanced samples are available at https://github.com/zhibinQiu/SRTNet.git.

Keywords

Cite

@article{arxiv.2210.16805,
  title  = {SRTNet: Time Domain Speech Enhancement Via Stochastic Refinement},
  author = {Zhibin Qiu and Mengfan Fu and Yinfeng Yu and LiLi Yin and Fuchun Sun and Hao Huang},
  journal= {arXiv preprint arXiv:2210.16805},
  year   = {2022}
}
R2 v1 2026-06-28T04:47:24.425Z