English
Related papers

Related papers: FoleyDirector: Fine-Grained Temporal Steering for …

200 papers

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporally aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Mengtian Li , Kunyan Dai , Yi Ding , Ruobing Ni , Ying Zhang , Wenwu Wang , Zhifeng Xie

Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) generation continues to be a challenge, principally because of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-20 Xinhao Mei , Varun Nagaraja , Gael Le Lan , Zhaoheng Ni , Ernie Chang , Yangyang Shi , Vikas Chandra

Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions are recreated to match the on-screen motion. This process is…

The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal…

Sound · Computer Science 2024-12-25 Yaoyun Zhang , Xuenan Xu , Mengyue Wu

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability…

Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono audio generation that lacks spatial perception, while the…

Sound · Computer Science 2025-08-22 Lei Zhao , Rujin Chen , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounter limitations when…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yiming Zhang , Yicheng Gu , Yanhong Zeng , Zhening Xing , Yuancheng Wang , Zhizheng Wu , Kai Chen

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents…

Sound · Computer Science 2024-09-12 Qi Yang , Binjie Mao , Zili Wang , Xing Nie , Pengfei Gao , Ying Guo , Cheng Zhen , Pengfei Yan , Shiming Xiang

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires…

Sound · Computer Science 2025-11-25 Satvik Dixit , Koichi Saito , Zhi Zhong , Yuki Mitsufuji , Chris Donahue

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shentong Mo , Yibing Song

We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM…

Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound…

Sound · Computer Science 2025-09-18 Junwon Lee , Jaekwon Im , Dabin Kim , Juhan Nam

Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research…

Sound · Computer Science 2025-07-21 Zhi Zhong , Akira Takahashi , Shuyang Cui , Keisuke Toyama , Shusuke Takahashi , Yuki Mitsufuji

Foley Control is a lightweight approach to video-guided Foley that keeps pretrained single-modality models frozen and learns only a small cross-attention bridge between them. We connect V-JEPA2 video embeddings to a frozen Stable Audio Open…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Ciara Rowles , Varun Jampani , Simon Donné , Shimon Vainer , Julian Parker , Zach Evans

Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language, a critical component of artistic expression in filmmaking.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Feizhen Huang , Yu Wu , Yutian Lin , Bo Du

Foley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advancements in deep…

Sound · Computer Science 2024-01-18 Yoonjin Chung , Junwon Lee , Juhan Nam

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating…

Sound · Computer Science 2026-03-23 Pengjun Fang , Yingqing He , Yazhou Xing , Qifeng Chen , Ser-Nam Lim , Harry Yang

Video-to-audio (V2A) generation utilizes visual-only video features to produce realistic sounds that correspond to the scene. However, current V2A models often lack fine-grained control over the generated audio, especially in terms of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Bingliang Li , Fengyu Yang , Yuxin Mao , Qingwen Ye , Hongkai Chen , Yiran Zhong
‹ Prev 1 2 3 10 Next ›