中文
相关论文

相关论文: Scaling Up Audio-Synchronized Visual Animation: An…

200 篇论文

Aligning video sequences is a fundamental yet still unsolved component for a broad range of applications in computer graphics and vision. Most classical image processing methods cannot be directly applied to related video problems due to…

计算机视觉与模式识别 · 计算机科学 2017-09-19 Patrick Wieschollek , Ido Freeman , Hendrik P. A. Lensch

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Xinlong Chen , Yue Ding , Weihong Lin , Jingyun Hua , Linli Yao , Yang Shi , Bozhou Li , Yuanxing Zhang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

Integrating audio and visual data for training multimodal foundational models remains a challenge. The Audio-Video Vector Alignment (AVVA) framework addresses this by considering AV scene alignment beyond mere temporal synchronization, and…

多媒体 · 计算机科学 2025-11-12 Ali Vosoughi , Dimitra Emmanouilidou , Hannes Gamper

We propose a self-supervised learning approach for videos that learns representations of both the RGB frames and the accompanying audio without human supervision. In contrast to images that capture the static scene appearance, videos also…

计算机视觉与模式识别 · 计算机科学 2023-02-16 Simon Jenni , Alexander Black , John Collomosse

We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling…

声音 · 计算机科学 2025-10-06 Amir Dellali , Luca A. Lanzendörfer , Florian Grötschla , Roger Wattenhofer

The scarcity of labeled audio-visual datasets is a constraint for training superior audio-visual speaker diarization systems. To improve the performance of audio-visual speaker diarization, we leverage pre-trained supervised and…

音频与语音处理 · 电气工程与系统科学 2023-12-08 Huan Zhao , Li Zhang , Yue Li , Yannan Wang , Hongji Wang , Wei Rao , Qing Wang , Lei Xie

The practicality of a video surveillance system is adversely limited by the amount of queries that can be placed on human resources and their vigilance in response. To transcend this limitation, a major effort under way is to include…

计算机视觉与模式识别 · 计算机科学 2014-05-16 Samaneh Khoshrou , Jaime S. Cardoso , Luis F. Teixeira

High-quality AI-powered video dubbing demands precise audio-lip synchronization, high-fidelity visual generation, and faithful preservation of identity and background. Most existing methods rely on a mask-based training strategy, where the…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Xindi Zhang , Dechao Meng , Steven Xiao , Qi Wang , Peng Zhang , Bang Zhang

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evolution. In this work, we…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Bohan Li , Shuojue Yang , Baorui Peng , Xianda Guo , Erli Zhang , Youqi Tao , Junfeng Duan , Daguang Xu , Qi Dou , Xin Jin , Wenjun Zeng , Hao Zhao , Yueming Jin

Despite that convolution neural networks (CNN) have recently demonstrated high-quality reconstruction for video super-resolution (VSR), efficiently training competitive VSR models remains a challenging problem. It usually takes an order of…

计算机视觉与模式识别 · 计算机科学 2022-05-18 Lijian Lin , Xintao Wang , Zhongang Qi , Ying Shan

An audio-visual event (AVE) is denoted by the correspondence of the visual and auditory signals in a video segment. Precise localization of the AVEs is very challenging since it demands effective multi-modal feature correspondence to ground…

计算机视觉与模式识别 · 计算机科学 2022-10-12 Tanvir Mahmud , Diana Marculescu

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yehang Zhang , Xinli Xu , Xiaojie Xu , Li Liu , Yingcong Chen

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

音频与语音处理 · 电气工程与系统科学 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

We study self-supervised video representation learning, which is a challenging task due to 1) lack of labels for explicit supervision; 2) unstructured and noisy visual information. Existing methods mainly use contrastive loss with video…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Deng Huang , Wenhao Wu , Weiwen Hu , Xu Liu , Dongliang He , Zhihua Wu , Xiangmiao Wu , Mingkui Tan , Errui Ding

Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Xinyu Zhang , Dong Gong , Zicheng Duan , Anton van den Hengel , Lingqiao Liu

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yikang Ding , Jiwen Liu , Wenyuan Zhang , Zekun Wang , Wentao Hu , Liyuan Cui , Mingming Lao , Yingchao Shao , Hui Liu , Xiaohan Li , Ming Chen , Xiaoqiang Liu , Yu-Shen Liu , Pengfei Wan

Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: Does audio-video joint denoising training improve video…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Jianzong Wu , Hao Lian , Dachao Hao , Ye Tian , Qingyu Shi , Biaolong Chen , Hao Jiang , Yunhai Tong

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

计算机视觉与模式识别 · 计算机科学 2019-08-22 Ruohan Gao , Kristen Grauman

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localization without…

计算机视觉与模式识别 · 计算机科学 2021-12-23 Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , Ji-Rong Wen

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…