English
Related papers

Related papers: SyncFusion: Multimodal Onset-synchronized Video-to…

200 papers

Foley sound presents the background sound for multimedia content and the generation of Foley sound involves computationally modelling sound effects with specialized techniques. In this work, we proposed a system for DCASE 2023 challenge…

Sound · Computer Science 2023-09-18 Yi Yuan , Haohe Liu , Xubo Liu , Xiyuan Kang , Mark D. Plumbley , Wenwu Wang

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2)…

Sound · Computer Science 2025-11-13 Shulei Ji , Zihao Wang , Jiaxing Yu , Xiangyuan Yang , Shuyu Li , Songruoyao Wu , Kejun Zhang

With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Haonan Qiu , Menghan Xia , Yong Zhang , Yingqing He , Xintao Wang , Ying Shan , Ziwei Liu

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

Sound · Computer Science 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Kaixin Shen , Ruijie Quan , Linchao Zhu , Jun Xiao , Yi Yang

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

Computer Vision and Pattern Recognition · Computer Science 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an…

Computer Vision and Pattern Recognition · Computer Science 2020-11-10 Kun Su , Xiulong Liu , Eli Shlizerman

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective,…

In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM…

Providing soundtracks for videos remains a costly and time-consuming challenge for multimedia content creators. We introduce EMSYNC, an automatic video-based symbolic music generator that creates music aligned with a video's emotional…

Sound · Computer Science 2026-02-06 Serkan Sulun , Paula Viana , Matthew E. P. Davies

Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality. In response to these limitations, we introduce MoFusion, i.e., a new…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Rishabh Dabral , Muhammad Hamza Mughal , Vladislav Golyanik , Christian Theobalt

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, \ie synthesizing missing audio segments that correspond to their accompanying…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Hang Zhou , Ziwei Liu , Xudong Xu , Ping Luo , Xiaogang Wang

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Huadai Liu , Jialei Wang , Rongjie Huang , Yang Liu , Heng Lu , Zhou Zhao , Wei Xue

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

Computer Vision and Pattern Recognition · Computer Science 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

The world of audio production and design has long been a difficult one to break into, requiring expertise and a working knowledge of the standard digital audio paradigms. This paper describes a novel interface that makes audio production…

Human-Computer Interaction · Computer Science 2020-10-01 Alexander Scarlatos

In video game design, audio (both environmental background music and object sound effects) play a critical role. Sounds are typically pre-created assets designed for specific locations or objects in a game. However, user-generated content…

Human-Computer Interaction · Computer Science 2024-04-29 Thomas Marrinan , Pakeeza Akram , Oli Gurmessa , Anthony Shishkin

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Yaniv Nikankin , Niv Haim , Michal Irani

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li
‹ Prev 1 4 5 6 7 8 10 Next ›