中文

基于扩散的生成式方法与判别式方法在语音修复中的分析比较

音频与语音处理 2023-03-17 v2 机器学习 声音

摘要

近年来,基于扩散的生成模型对计算机视觉与语音处理领域产生了重大影响。除数据生成任务外,它们也被用于语音增强与去混响等数据修复任务。虽然传统上认为判别式模型更强大(例如在语音增强中),但近期研究表明生成式扩散方法已大幅缩小了这一性能差距。本文中,我们系统比较了生成式扩散模型与判别式方法在不同语音修复任务上的性能。为此,我们将先前在复数时频域基于扩散的语音增强工作扩展至带宽扩展任务。随后,我们在语音去噪、去混响和带宽扩展三项修复任务上,将其与具有相同网络架构的判别训练神经网络进行比较。我们观察到,生成式方法在所有任务上总体优于其判别式对应方法,且在非加性失真模型(如去混响与带宽扩展)中获益最为显著。代码与音频示例可在 https://uhh.de/inf-sp-sgmsemultitask 在线获取。

关键词

引用

@article{arxiv.2211.02397,
  title  = {Analysing Diffusion-based Generative Approaches versus Discriminative Approaches for Speech Restoration},
  author = {Jean-Marie Lemercier and Julius Richter and Simon Welker and Timo Gerkmann},
  journal= {arXiv preprint arXiv:2211.02397},
  year   = {2023}
}

备注

\c{opyright} 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works