English
Related papers

Related papers: AVI-Edit: Audio-sync Video Instance Editing with G…

200 papers

Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Amir Mohammad Izadi , Seyed Mohammad Hadi Hosseini , Soroush Vafaie Tabar , Ali Abdollahi , Armin Saghafian , Mahdieh Soleymani Baghshah

Video editing according to instructions is a highly challenging task due to the difficulty in collecting large-scale, high-quality edited video pair data. This scarcity not only limits the availability of training data but also hinders the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Chi Zhang , Chengjian Feng , Feng Yan , Qiming Zhang , Mingjin Zhang , Yujie Zhong , Jing Zhang , Lin Ma

Modern one-stage video instance segmentation networks suffer from two limitations. First, convolutional features are neither aligned with anchor boxes nor with ground-truth bounding boxes, reducing the mask sensitivity to spatial location.…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 Minghan Li , Shuai Li , Lida Li , Lei Zhang

Recent research has witnessed the advances in facial image editing tasks. For video editing, however, previous methods either simply apply transformations frame by frame or utilize multiple frames in a concatenated or iterative fashion,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-06 Meng Cao , Haozhi Huang , Hao Wang , Xuan Wang , Li Shen , Sheng Wang , Linchao Bao , Zhifeng Li , Jiebo Luo

Open-vocabulary image segmentation has been advanced through the synergy between mask generators and vision-language models like Contrastive Language-Image Pre-training (CLIP). Previous approaches focus on generating masks while aligning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Quan-Sheng Zeng , Yunheng Li , Daquan Zhou , Guanbin Li , Qibin Hou , Ming-Ming Cheng

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Inkyu Shin , Qihang Yu , Xiaohui Shen , In So Kweon , Kuk-Jin Yoon , Liang-Chieh Chen

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior…

Sound · Computer Science 2023-08-02 Chen Liu , Peike Li , Xingqun Qi , Hu Zhang , Lincheng Li , Dadong Wang , Xin Yu

In the art of video editing, sound helps add character to an object and immerse the viewer within a space. Through formative interviews with professional editors (N=10), we found that the task of adding sounds to video can be challenging.…

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Xiaodi Li , Pan Xie , Yi Ren , Qijun Gan , Chen Zhang , Fangyuan Kong , Xiang Yin , Bingyue Peng , Zehuan Yuan

The goal of this paper is to discover, segment, and track independently moving objects in complex visual scenes. Previous approaches have explored the use of optical flow for motion segmentation, leading to imperfect predictions due to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Junyu Xie , Weidi Xie , Andrew Zisserman

In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yan-Bo Lin , Kevin Lin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Chung-Ching Lin , Xiaofei Wang , Gedas Bertasius , Lijuan Wang

Most real-world image editing tasks require multiple sequential edits to achieve desired results. Current editing approaches, primarily designed for single-object modifications, struggle with sequential editing: especially with maintaining…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Daneul Kim , Jaeah Lee , Jaesik Park

Instruction-based video editing requires transforming a source video according to a natural-language instruction while preserving irrelevant content and remaining temporally coherent. We argue that existing Diffusion Transformer (DiT)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yan Li , Lin Liu , Xiaopeng Zhang , Qi Tian

Image generation and editing have seen a great deal of advancements with the rise of large-scale diffusion models that allow user control of different modalities such as text, mask, depth maps, etc. However, controlled editing of videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 AmirHossein Zamani , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 George Sterpu , Christian Saam , Naomi Harte

Autoregressive video diffusion models (AR-VDMs) show strong promise as scalable alternatives to bidirectional VDMs, enabling real-time and interactive applications. Yet there remains room for improvement in their sample fidelity. A…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zhengyang Yu , Akio Hayakawa , Masato Ishii , Qingtao Yu , Takashi Shibuya , Jing Zhang , Yuki Mitsufuji

Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural…

Sound · Computer Science 2024-07-16 Santiago Pascual , Chunghsin Yeh , Ioannis Tsiamas , Joan Serrà

Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Huanghao Yin , Shenkun Xu , Kanle Shi , Junhai Yong , Bin Wang

Real-world image matting is essential for applications in content creation and augmented reality. However, it remains challenging due to the complex nature of scenes and the scarcity of high-quality datasets. To address these limitations,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Rui Liu

As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-world online videos are often plagued by degradations such as blurring and quantization noise, due to the high…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Yutong Wang , Jiajie Teng , Jiajiong Cao , Yuming Li , Chenguang Ma , Hongteng Xu , Dixin Luo