English

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Sound 2024-12-13 v1 Computer Vision and Pattern Recognition Multimedia Audio and Speech Processing

Abstract

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled data in real-world scenes, we introduce YingSound, a foundation model designed for video-guided sound generation that supports high-quality audio generation in few-shot settings. Specifically, YingSound consists of two major modules. The first module uses a conditional flow matching transformer to achieve effective semantic alignment in sound generation across audio and visual modalities. This module aims to build a learnable audio-visual aggregator (AVA) that integrates high-resolution visual features with corresponding audio features at multiple stages. The second module is developed with a proposed multi-modal visual-audio chain-of-thought (CoT) approach to generate finer sound effects in few-shot settings. Finally, an industry-standard video-to-audio (V2A) dataset that encompasses various real-world scenarios is presented. We show that YingSound effectively generates high-quality synchronized sounds across diverse conditional inputs through automated evaluations and human studies. Project Page: \url{https://giantailab.github.io/yingsound/}

Keywords

Cite

@article{arxiv.2412.09168,
  title  = {YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls},
  author = {Zihao Chen and Haomin Zhang and Xinhan Di and Haoyu Wang and Sizhe Shan and Junjie Zheng and Yunming Liang and Yihan Fan and Xinfa Zhu and Wenjie Tian and Yihua Wang and Chaofan Ding and Lei Xie},
  journal= {arXiv preprint arXiv:2412.09168},
  year   = {2024}
}

Comments

16 pages, 4 figures