中文
相关论文

相关论文: Masked Generative Video-to-Audio Transformers with…

200 篇论文

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then…

多媒体 · 计算机科学 2026-03-18 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

音频与语音处理 · 电气工程与系统科学 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual-acoustic content. However, existing open-source methods mainly rely on either dual-tower designs with posterior alignment or fully…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Longbin Ji , Guan Wang , Xuan Wei , Chenye Yang , Xiangrui Liu , Zhenyu Zhang , Shuohuan Wang , Yu Sun , Jingzhou He

Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a…

声音 · 计算机科学 2025-01-07 Yongqi Wang , Wenxiang Guo , Rongjie Huang , Jiawei Huang , Zehan Wang , Fuming You , Ruiqi Li , Zhou Zhao

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no model has yet led…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Xusen Sun , Longhao Zhang , Hao Zhu , Peng Zhang , Bang Zhang , Xinya Ji , Kangneng Zhou , Daiheng Gao , Liefeng Bo , Xun Cao

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yan-Bo Lin , Jonah Casebeer , Long Mai , Aniruddha Mahapatra , Gedas Bertasius , Nicholas J. Bryan

Articulatory-to-acoustic (A2A) synthesis refers to the generation of audible speech from captured movement of the speech articulators. This technique has numerous applications, such as restoring oral communication to people who cannot…

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Guozhen Zhang , Zixiang Zhou , Teng Hu , Ziqiao Peng , Youliang Zhang , Yi Chen , Yuan Zhou , Qinglin Lu , Limin Wang

Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Guangzhao Li , Kangrui Cen , Baixuan Zhao , Yi Xin , Siqi Luo , Guangtao Zhai , Lei Zhang , Xiaohong Liu

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yongtao Ge , Kangyang Xie , Guangkai Xu , Mingyu Liu , Li Ke , Longtao Huang , Hui Xue , Hao Chen , Chunhua Shen

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

计算机视觉与模式识别 · 计算机科学 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce…

Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box…

多媒体 · 计算机科学 2025-07-29 Junxian Wu , Weitao You , Heda Zuo , Dengming Zhang , Pei Chen , Lingyun Sun

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos…

多媒体 · 计算机科学 2024-09-12 Yan-Bo Lin , Yu Tian , Linjie Yang , Gedas Bertasius , Heng Wang

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

多媒体 · 计算机科学 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more challenging setting:…

声音 · 计算机科学 2025-07-28 Fang Kang , Yin Cao , Haoyu Chen

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

计算机视觉与模式识别 · 计算机科学 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

声音 · 计算机科学 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid representations and…

图形学 · 计算机科学 2023-06-21 Yue Yang , Kaipeng Zhang , Yuying Ge , Wenqi Shao , Zeyue Xue , Yu Qiao , Ping Luo