English
Related papers

Related papers: Video-Guided Foley Sound Generation with Multimoda…

200 papers

Existing works have made strides in video generation, but the lack of sound effects (SFX) and background music (BGM) hinders a complete and immersive viewer experience. We introduce a novel semantically consistent v ideo-to-audio generation…

Multimedia · Computer Science 2024-04-29 Gehui Chen , Guan'an Wang , Xiaowen Huang , Jitao Sang

AI creation, such as poem or lyrics generation, has attracted increasing attention from both industry and academic communities, with many promising models proposed in the past few years. Existing methods usually estimate the outputs based…

Artificial Intelligence · Computer Science 2024-09-05 Qian Cao , Xu Chen , Ruihua Song , Hao Jiang , Guang Yang , Zhao Cao

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and…

Multimedia · Computer Science 2025-11-27 Xinyue Guo , Xiaoran Yang , Lipan Zhang , Jianxuan Yang , Zhao Wang , Jian Luan

Visuals can enhance our experience of music, owing to the way they can amplify the emotions and messages conveyed within it. However, creating music visualization is a complex, time-consuming, and resource-intensive process. We introduce…

Human-Computer Interaction · Computer Science 2023-09-29 Vivian Liu , Tao Long , Nathan Raw , Lydia Chilton

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation…

Sound · Computer Science 2023-07-03 Simian Luo , Chuanhao Yan , Chenxu Hu , Hang Zhao

Video inbetweening creates smooth and natural transitions between two image frames, making it an indispensable tool for video editing and long-form video synthesis. Existing works in this domain are unable to generate large, complex, or…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao

Generating long and consistent videos has emerged as a significant yet challenging problem. While most existing diffusion-based video generation models, derived from image generation models, demonstrate promising performance in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yichen Ouyang , jianhao Yuan , Hao Zhao , Gaoang Wang , Bo zhao

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporally aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Mengtian Li , Kunyan Dai , Yi Ding , Ruobing Ni , Ying Zhang , Wenwu Wang , Zhifeng Xie

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired…

Sound · Computer Science 2024-10-08 Han Yang , Kun Su , Yutong Zhang , Jiaben Chen , Kaizhi Qian , Gaowen Liu , Chuang Gan

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating…

Sound · Computer Science 2024-03-21 Ronghui Li , Yuqin Dai , Yachao Zhang , Jun Li , Jian Yang , Jie Guo , Xiu Li

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see (the video's image…

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Qiang Zhou , Shaofeng Zhang , Nianzu Yang , Ye Qian , Hao Li

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…

Multimedia · Computer Science 2022-11-22 Chenye Cui , Yi Ren , Jinglin Liu , Rongjie Huang , Zhou Zhao

We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is achieved by combining two distinct models. The first model,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-13 Kevin Kilgour , Beat Gfeller , Qingqing Huang , Aren Jansen , Scott Wisdom , Marco Tagliasacchi

Recent advancements in video generation models have significantly improved their ability to follow text prompts. However, the customization of dynamic visual effects, defined as temporally evolving and appearance-driven visual phenomena…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Rui Zhao , Mike Zheng Shou

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

Multimedia · Computer Science 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be described through text prompts or image references. Given a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Leheng Li , Weichao Qiu , Xu Yan , Jing He , Kaiqiang Zhou , Yingjie Cai , Qing Lian , Bingbing Liu , Ying-Cong Chen