English
Related papers

Related papers: SpongeBob: Sync-Aware Harmonious Audio-Visual Gene…

200 papers

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yaru Chen , Ruohao Guo , Liting Gao , Yang Xiang , Qingyu Luo , Zhenbo Li , Wenwu Wang

Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-08 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Xinyu Zhang , Dong Gong , Zicheng Duan , Anton van den Hengel , Lingqiao Liu

Given a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track. Existing state-of-the-art V2C models…

Computation and Language · Computer Science 2024-07-03 Gaoxiang Cong , Yuankai Qi , Liang Li , Amin Beheshti , Zhedong Zhang , Anton van den Hengel , Ming-Hsuan Yang , Chenggang Yan , Qingming Huang

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

Sound · Computer Science 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hongye Cheng , Tianyu Wang , Guangsi Shi , Zexing Zhao , Yanwei Fu

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robustness to noisy…

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Aashish Rai , Srinath Sridhar

This presentation introduces a self-supervised learning approach to the synthesis of new video clips from old ones, with several new key elements for improved spatial resolution and realism: It conditions the synthesis process on contextual…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Guillaume Le Moing , Jean Ponce , Cordelia Schmid

There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with object-level audio-visual operations. Specifically,…

Multimedia · Computer Science 2025-10-02 Youquan Fu , Ruiyang Si , Hongfa Wang , Dongzhan Zhou , Jiacheng Sun , Ping Luo , Di Hu , Hongyuan Zhang , Xuelong Li

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-stage pipeline that first generates mono audio and then…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Mengchen Zhang , Qi Chen , Tong Wu , Zihan Liu , Dahua Lin

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-23 Chenyu Yang , Shuai Wang , Hangting Chen , Wei Tan , Jianwei Yu , Haizhou Li

Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods often face temporal misalignment, where audio cues and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Kexin Li , Zongxin Yang , Yi Yang , Jun Xiao

Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly…

Sound · Computer Science 2025-12-03 Xueyan Li , Yuxin Wang , Mengjie Jiang , Qingzi Zhu , Jiang Zhang , Zoey Kim , Yazhe Niu

We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio,…

Video editing serves as a fundamental pillar of digital media, spanning applications in entertainment, education, and professional communication. However, previous methods often overlook the necessity of comprehensively understanding both…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jing Gu , Yuwei Fang , Ivan Skorokhodov , Peter Wonka , Xinya Du , Sergey Tulyakov , Xin Eric Wang

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interplay between audio and…

Multimedia · Computer Science 2024-10-01 Kun Su , Xiulong Liu , Eli Shlizerman