SINGER:基于多尺度谱系扩散模型的 vivid 音频驱动唱歌视频生成
摘要
近期在 generative models 中的进展显著提升了 talking face video generation 的 performance,但 singing video generation 仍未得到充分探索。talking 和 singing 之间的差异限制了 existing talking face video generation models when applied to singing 的 performance。talking 和 singing-specifically 在 audio characteristics 和 behavioral expressions 方面的 fundamental differences limit effectiveness of existing models。我们观察到,singing 和 talking audio 的差异在 frequency 和 amplitude 方面 manifest。为 address this,我们设计了 multi-scale spectral module help model learn singing patterns in spectral domain。此外,我们 develop a spectral-filtering module aids model learn 人类与 singing audio 相关的 behaviors。这两个模块集成到 diffusion model 中 to enhance singing video generation performance,resulting in our proposed model SINGER。此外,lack of high-quality real-world singing face videos has hindered development of singing video generation community。为 address this gap,我们收集了一个 in-the-wild audio-visual singing dataset facilitate research in this area。我们的 experiments demonstrate that SINGER is capable of generating vivid singing videos and outperforms state-of-the-art methods in both objective and subjective evaluations。
引用
@article{arxiv.2412.03430,
title = {SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model},
author = {Yan Li and Ziya Zhou and Zhiqiang Wang and Wei Xue and Wenhan Luo and Yike Guo},
journal= {arXiv preprint arXiv:2412.03430},
year = {2024}
}