English
Related papers

Related papers: FoleyGRAM: Video-to-Audio Generation with GRAM-Ali…

200 papers

Recent advancements in audio generation have been spurred by the evolution of large-scale deep learning models and expansive datasets. However, the task of video-to-audio (V2A) generation continues to be a challenge, principally because of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-20 Xinhao Mei , Varun Nagaraja , Gael Le Lan , Zhaoheng Ni , Ernie Chang , Yangyang Shi , Vikas Chandra

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames. Previous studies…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Shentong Mo , Yibing Song

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class…

Machine Learning · Computer Science 2021-07-21 Sanchita Ghose , John J. Prevost

We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounter limitations when…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yiming Zhang , Yicheng Gu , Yanhong Zeng , Zhening Xing , Yuancheng Wang , Zhizheng Wu , Kai Chen

Recently, with the advancement of AIGC, deep learning-based video-to-audio (V2A) technology has garnered significant attention. However, existing research mostly focuses on mono audio generation that lacks spatial perception, while the…

Sound · Computer Science 2025-08-22 Lei Zhao , Rujin Chen , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to…

Sound · Computer Science 2025-03-12 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Rilin Chen , Yu Gu , Wei Liang , Dong Yu

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal…

Sound · Computer Science 2024-09-16 Zhiqi Huang , Dan Luo , Jun Wang , Huan Liao , Zhiheng Li , Zhiyong Wu

Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are…

Sound · Computer Science 2026-04-21 You Li , Dewei Zhou , Fan Ma , Fu Li , Dongliang He , Yi Yang

Traditional sound design workflows rely on manual alignment of audio events to visual cues, as in Foley sound design, where everyday actions like footsteps or object interactions are recreated to match the on-screen motion. This process is…

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporally aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Mengtian Li , Kunyan Dai , Yi Ding , Ruobing Ni , Ying Zhang , Wenwu Wang , Zhifeng Xie

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires…

Sound · Computer Science 2025-11-25 Satvik Dixit , Koichi Saito , Zhi Zhong , Yuki Mitsufuji , Chris Donahue

We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal…

Multimedia · Computer Science 2025-07-28 Hyunwoo Oh , SeungJu Cha , Kwanyoung Lee , Si-Woo Kim , Dong-Jin Kim

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley task, presents…

Sound · Computer Science 2024-09-12 Qi Yang , Binjie Mao , Zili Wang , Xing Nie , Pengfei Gao , Ying Guo , Cheng Zhen , Pengfei Yan , Shiming Xiang

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Yuchen Hu , Yu Gu , Chenxing Li , Rilin Chen , Dong Yu

Recent advances in video generation have achieved remarkable improvements in visual content fidelity. However, the absence of synchronized audio severely undermines immersive experience and restricts practical applications of these…

Sound · Computer Science 2025-12-09 Fu Li , Weichao Zhao , You Li , Zhichao Zhou , Dongliang He

Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound…

Sound · Computer Science 2025-09-18 Junwon Lee , Jaekwon Im , Dabin Kim , Juhan Nam

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Foley sound synthesis refers to the creation of authentic, diegetic sound effects for media, such as film or radio. In this study, we construct a neural Foley synthesizer capable of generating mono-audio clips across seven predefined…

Sound · Computer Science 2023-09-12 Ashwin Pillay , Sage Betko , Ari Liloia , Hao Chen , Ankit Shah

This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook…

Sound · Computer Science 2026-02-05 Kazuki Shimada , Christian Simon , Takashi Shibuya , Shusuke Takahashi , Yuki Mitsufuji

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Jiawei Liu , Weining Wang , Sihan Chen , Xinxin Zhu , Jing Liu
‹ Prev 1 2 3 10 Next ›