中文

利用StyleGAN2与梅尔谱生成多样化的情感发声爆发

声音 2022-06-28 v1 机器学习 音频与语音处理

摘要

我们描述了用于ICML表达性发声竞赛中生成式情感发声爆发任务(ExVo Generate)的方法。我们在预处理后音频样本的梅尔谱(mel-spectrogram)上训练了一个条件式StyleGAN2架构。模型生成的梅尔谱随后被逆变换回音频域。结果表明,对于所有情感,我们生成的样本在定性与定量角度均较竞赛提供的基线有实质性提升。更确切地说,即便对于我们表现最差的情感(敬畏,awe),我们也取得了1.76的FAD,而基线为4.81(作为参考,敬畏情感训练集与验证集之间的FAD为0.776)。

关键词

引用

@article{arxiv.2206.12563,
  title  = {Generating Diverse Vocal Bursts with StyleGAN2 and MEL-Spectrograms},
  author = {Marco Jiralerspong and Gauthier Gidel},
  journal= {arXiv preprint arXiv:2206.12563},
  year   = {2022}
}

备注

To be published at the ICML Expressive Vocalizations Workshop and Competition (ExVo Generate) held in conjunction with the 39th International Conference on Machine Learning