Autovocoder:利用可微数字信号处理从习得语音表征快速生成波形
声音
2023-05-25 v2 音频与语音处理
摘要
大多数最先进的文本到语音系统使用梅尔谱图作为中间表征,将任务分解为声学建模与波形生成。梅尔谱图通过简单、快速的 DSP 操作从波形提取,但从梅尔谱图生成高质量波形需要计算昂贵的机器学习:神经声码器。我们提出的“autovocoder”颠倒了这一安排。我们使用机器学习获得一种替代梅尔谱图的表征,该表征可使用包括可微逆 STFT 实现在内的简单、快速操作反变换回波形。Autovocoder 生成波形比基于 DSP 的 Griffin-Lim 算法快 5 倍,比神经声码器 HiFi-GAN 快 14 倍。我们提供感知听测结果以确认在复制合成任务中语音质量与 HiFi-GAN 相当。
引用
@article{arxiv.2211.06989,
title = {Autovocoder: Fast Waveform Generation from a Learned Speech Representation using Differentiable Digital Signal Processing},
author = {Jacob J Webber and Cassia Valentini-Botinhao and Evelyn Williams and Gustav Eje Henter and Simon King},
journal= {arXiv preprint arXiv:2211.06989},
year = {2023}
}
备注
Accepted to the 2023 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2023)