English
Related papers

Related papers: AVI-Edit: Audio-sync Video Instance Editing with G…

200 papers

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

Sound · Computer Science 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

Sound · Computer Science 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserving motion, temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 John Flynn , Wolfgang Paier , Dimitar Dinev , Sam Nhut Nguyen , Hayk Poghosyan , Manuel Toribio , Sandipan Banerjee , Guy Gafni

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Alejandro Pardo , Jui-Hsien Wang , Bernard Ghanem , Josef Sivic , Bryan Russell , Fabian Caba Heilbron

Parameter-efficient adaptation of vision-language foundation models is crucial for precise multimodal understanding of biomedical images, yet existing methods remain deterministic and often struggle under domain shift or ambiguous…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Taha Koleilat , Hassan Rivaz , Yiming Xiao

Eliminating time-consuming post-production processes and delivering high-quality videos in today's fast-paced digital landscape are the key advantages of real-time approaches. To address these needs, we present Real Time GAZED: a real-time…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Sudheer Achary , Rohit Girmaji , Adhiraj Anil Deshmukh , Vineet Gandhi

Large-scale generative models have achieved remarkable success in a number of domains. However, for sequential decision-making problems, such as robotics, action-labelled data is often scarce and therefore scaling-up foundation models for…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Marc Rigter , Tarun Gupta , Agrin Hilmkil , Chao Ma

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

AI-generated video has revolutionized short video production, filmmaking, and personalized media, making video local editing an essential tool. However, this progress also blurs the line between reality and fiction, posing challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Xuanyu Zhang , Youmin Xu , Runyi Li , Jiwen Yu , Weiqi Li , Zhipei Xu , Jian Zhang

We present Masked Audio-Video Learners (MAViL) to train audio-visual representations. Our approach learns with three complementary forms of self-supervision: (1) reconstruction of masked audio and video input data, (2) intra- and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Po-Yao Huang , Vasu Sharma , Hu Xu , Chaitanya Ryali , Haoqi Fan , Yanghao Li , Shang-Wen Li , Gargi Ghosh , Jitendra Malik , Christoph Feichtenhofer

Video Instance Segmentation (VIS) aims to simultaneously classify, segment, and track multiple object instances in videos. Recent clip-level VIS takes a short video clip as input each time showing stronger performance than frame-level VIS…

Computer Vision and Pattern Recognition · Computer Science 2022-03-04 Jialian Wu , Sudhir Yarram , Hui Liang , Tian Lan , Junsong Yuan , Jayan Eledath , Gerard Medioni

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

Sound · Computer Science 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar

The proliferation of video content production has led to vast amounts of data, posing substantial challenges in terms of analysis efficiency and resource utilization. Addressing this issue calls for the development of robust video analysis…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Ulindu De Silva , Leon Fernando , Kalinga Bandara , Rashmika Nawaratne

Video instance segmentation (VIS) is a new and critical task in computer vision. To date, top-performing VIS methods extend the two-stage Mask R-CNN by adding a tracking branch, leaving plenty of room for improvement. In contrast, we…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Dongfang Liu , Yiming Cui , Wenbo Tan , Yingjie Chen

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Tao Feng , Yifan Xie , Xun Guan , Jiyuan Song , Zhou Liu , Fei Ma , Fei Yu

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdles in (a) accurately…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Hyeonho Jeong , Geon Yeong Park , Jong Chul Ye

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we introduce VEGGIE, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Shoubin Yu , Difan Liu , Ziqiao Ma , Yicong Hong , Yang Zhou , Hao Tan , Joyce Chai , Mohit Bansal

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employing 2D generation or editing models to process each…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Liyi Chen , Ruihuang Li , Guowen Zhang , Pengfei Wang , Lei Zhang

Lip synchronization aims to generate realistic talking videos that match given audio, which is essential for high-quality video dubbing. However, current methods have fundamental drawbacks: mask-based approaches suffer from local color…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Ruidi Fan , Yang Zhou , Siyuan Wang , Tian Yu , Yutong Jiang , Xusheng Liu
‹ Prev 1 8 9 10 Next ›