中文

Autovocoder:利用可微数字信号处理从习得语音表征快速生成波形

声音 2023-05-25 v2 音频与语音处理

摘要

大多数最先进的文本到语音系统使用梅尔谱图作为中间表征,将任务分解为声学建模与波形生成。梅尔谱图通过简单、快速的 DSP 操作从波形提取,但从梅尔谱图生成高质量波形需要计算昂贵的机器学习:神经声码器。我们提出的“autovocoder”颠倒了这一安排。我们使用机器学习获得一种替代梅尔谱图的表征,该表征可使用包括可微逆 STFT 实现在内的简单、快速操作反变换回波形。Autovocoder 生成波形比基于 DSP 的 Griffin-Lim 算法快 5 倍,比神经声码器 HiFi-GAN 快 14 倍。我们提供感知听测结果以确认在复制合成任务中语音质量与 HiFi-GAN 相当。

关键词

引用

@article{arxiv.2211.06989,
  title  = {Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing},
  author = {Jacob J Webber and Cassia Valentini-Botinhao and Evelyn Williams and Gustav Eje Henter and Simon King},
  journal= {arXiv preprint arXiv:2211.06989},
  year   = {2023}
}

备注

Accepted to the 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023)