中文
相关论文

相关论文: SyncFusion: Multimodal Onset-synchronized Video-to…

200 篇论文

Foley sound presents the background sound for multimedia content and the generation of Foley sound involves computationally modelling sound effects with specialized techniques. In this work, we proposed a system for DCASE 2023 challenge…

声音 · 计算机科学 2023-09-18 Yi Yuan , Haohe Liu , Xubo Liu , Xiyuan Kang , Mark D. Plumbley , Wenwu Wang

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2)…

声音 · 计算机科学 2025-11-13 Shulei Ji , Zihao Wang , Jiaxing Yu , Xiangyuan Yang , Shuyu Li , Songruoyao Wu , Kejun Zhang

With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of…

计算机视觉与模式识别 · 计算机科学 2024-01-31 Haonan Qiu , Menghan Xia , Yong Zhang , Yingqing He , Xintao Wang , Ying Shan , Ziwei Liu

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

声音 · 计算机科学 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Kaixin Shen , Ruijie Quan , Linchao Zhu , Jun Xiao , Yi Yang

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Kun Su , Xiulong Liu , Eli Shlizerman

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective,…

In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM…

Providing soundtracks for videos remains a costly and time-consuming challenge for multimedia content creators. We introduce EMSYNC, an automatic video-based symbolic music generator that creates music aligned with a video's emotional…

声音 · 计算机科学 2026-02-06 Serkan Sulun , Paula Viana , Matthew E. P. Davies

Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality. In response to these limitations, we introduce MoFusion, i.e., a new…

计算机视觉与模式识别 · 计算机科学 2023-05-16 Rishabh Dabral , Muhammad Hamza Mughal , Vladislav Golyanik , Christian Theobalt

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, \ie synthesizing missing audio segments that correspond to their accompanying…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Hang Zhou , Ziwei Liu , Xudong Xu , Ping Luo , Xiaogang Wang

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Huadai Liu , Jialei Wang , Rongjie Huang , Yang Liu , Heng Lu , Zhou Zhao , Wei Xue

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

音频与语音处理 · 电气工程与系统科学 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

计算机视觉与模式识别 · 计算机科学 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

The world of audio production and design has long been a difficult one to break into, requiring expertise and a working knowledge of the standard digital audio paradigms. This paper describes a novel interface that makes audio production…

人机交互 · 计算机科学 2020-10-01 Alexander Scarlatos

In video game design, audio (both environmental background music and object sound effects) play a critical role. Sounds are typically pre-created assets designed for specific locations or objects in a game. However, user-generated content…

人机交互 · 计算机科学 2024-04-29 Thomas Marrinan , Pakeeza Akram , Oli Gurmessa , Anthony Shishkin

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Yaniv Nikankin , Niv Haim , Michal Irani

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Junjia Huang , Binbin Yang , Pengxiang Yan , Jiyang Liu , Bin Xia , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li