中文
相关论文

相关论文: AV-Link: Temporally-Aligned Diffusion Features for…

200 篇论文

Audio-Visual Localization (AVL) aims to identify sound-emitting sources within a visual scene. However, existing studies focus on image-level audio-visual associations, failing to capture temporal dynamics. Moreover, they assume simplified…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Hahyeon Choi , Junhoo Lee , Nojun Kwak

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data, such as conflating…

声音 · 计算机科学 2026-03-23 Pengjun Fang , Yingqing He , Yazhou Xing , Qifeng Chen , Ser-Nam Lim , Harry Yang

We present a method to create diffusion-based video models from pretrained Text-to-Image (T2I) models. Recently, AnimateDiff proposed freezing the T2I model while only training temporal layers. We advance this method by proposing a unique…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Mingi Kwon , Seoung Wug Oh , Yang Zhou , Difan Liu , Joon-Young Lee , Haoran Cai , Baqiao Liu , Feng Liu , Youngjung Uh

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consistency due to limited natural language understanding and data…

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

With the increasing adoption of video anomaly detection in intelligent surveillance domains, conventional visual-based detection approaches often struggle with information insufficiency and high false-positive rates in complex environments.…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Peng Wu , Wanshun Su , Guansong Pang , Yujia Sun , Qingsen Yan , Peng Wang , Yanning Zhang

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Daili Hua , Xizhi Wang , Bohan Zeng , Xinyi Huang , Hao Liang , Junbo Niu , Xinlong Chen , Quanqing Xu , Wentao Zhang

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to…

声音 · 计算机科学 2025-03-12 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Rilin Chen , Yu Gu , Wei Liang , Dong Yu

AIGC has rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Kai Liu , Yanhao Zheng , Kai Wang , Shengqiong Wu , Rongjunchen Zhang , Jiebo Luo , Dimitrios Hatzinakos , Ziwei Liu , Hao Fei , Tat-Seng Chua

Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jiayu Zhang , Shuo Ye , Qilang Ye , Xun Lin , Zihan Song , Zitong Yu

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions…

Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Ye Tian , Ling Yang , Haotian Yang , Yuan Gao , Yufan Deng , Jingmin Chen , Xintao Wang , Zhaochen Yu , Xin Tao , Pengfei Wan , Di Zhang , Bin Cui

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Video Diffusion Models (VDMs) offer a promising approach for simulating dynamic scenes and environments, with broad applications in robotics and media generation. However, existing models often generate temporally incoherent content that…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhexiao Xiong , Yizhi Song , Liu He , Wei Xiong , Yu Yuan , Feng Qiao , Nathan Jacobs

This work addresses the lack of multimodal generative models capable of producing high-quality videos with spatially aligned audio. While recent advancements in generative models have been successful in video generation, they often overlook…

声音 · 计算机科学 2026-02-05 Kazuki Shimada , Christian Simon , Takashi Shibuya , Shusuke Takahashi , Yuki Mitsufuji