English
Related papers

Related papers: AVI-Edit: Audio-sync Video Instance Editing with G…

200 papers

Language-referred audio-visual segmentation (Ref-AVS) aims to segment target objects described by natural language by jointly reasoning over video, audio, and text. Beyond generating segmentation masks, providing rich and interpretable…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Jinxing Zhou , Yanghao Zhou , Yaoting Wang , Zongyan Han , Jiaqi Ma , Henghui Ding , Rao Muhammad Anwer , Hisham Cholakkal

In this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visual modeling can improve class-incremental learning, but…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Weiguo Pian , Shentong Mo , Yunhui Guo , Yapeng Tian

Facial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenges persist in achieving precise, 3D-aware attribute…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yu-Kai Huang , Yutong Zheng , Yen-Shuo Su , Anudeepsekhar Bolimera , Han Zhang , Fangyi Chen , Marios Savvides

We introduce VIVE3D, a novel approach that extends the capabilities of image-based 3D GANs to video editing and is able to represent the input video in an identity-preserving and temporally consistent way. We propose two new building…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Anna Frühstück , Nikolaos Sarafianos , Yuanlu Xu , Peter Wonka , Tony Tung

In this report, we present MagicEdit, a surprisingly simple yet effective solution to the text-guided video editing task. We found that high-fidelity and temporally coherent video-to-video translation can be achieved by explicitly…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Jun Hao Liew , Hanshu Yan , Jianfeng Zhang , Zhongcong Xu , Jiashi Feng

Current benchmarks for video segmentation are limited to annotating only salient objects (i.e., foreground instances). Despite their impressive architectural designs, previous works trained on these benchmarks have struggled to adapt to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Sangbeom Lim , Seongchan Kim , Seungjun An , Seokju Cho , Paul Hongsuck Seo , Seungryong Kim

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Daniel Gilo , Or Litany

Diffusion-based text-to-image (T2I) models have demonstrated remarkable results in global video editing tasks. However, their focus is primarily on global video modifications, and achieving desired attribute-specific changes remains a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Haoyu Zheng , Wenqiao Zhang , Zheqi Lv , Yu Zhong , Yang Dai , Jianxiang An , Yongliang Shen , Juncheng Li , Dongping Zhang , Siliang Tang , Yueting Zhuang

Labeling pixel-wise object masks in videos is a resource-intensive and laborious process. Box-supervised Video Instance Segmentation (VIS) methods have emerged as a viable solution to mitigate the labor-intensive annotation process. . In…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Zhangjing Yang , Dun Liu , Wensheng Cheng , Jinqiao Wang , Yi Wu

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Shuchen Weng , Haojie Zheng , Peixuan Zhang , Yuchen Hong , Han Jiang , Si Li , Boxin Shi

The explosive growth of video streaming presents challenges in achieving high accuracy and low training costs for video-language retrieval. However, existing methods rely on large-scale pre-training to improve video retrieval performance,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Haoyu Zhao , Jiaxi Gu , Shicong Wang , Xing Zhang , Hang Xu , Zuxuan Wu , Yu-Gang Jiang

Recent works have explored text-guided image editing using diffusion models and generated edited images based on text prompts. However, the models struggle to accurately locate the regions to be edited and faithfully perform precise edits.…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Qian Wang , Biao Zhang , Michael Birsak , Peter Wonka

This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remarkable ability to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Zhichao Zuo , Zhao Zhang , Yan Luo , Yang Zhao , Haijun Zhang , Yi Yang , Meng Wang

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

Text-to-image diffusion models have demonstrated remarkable progress in synthesizing high-quality images from text prompts, which boosts researches on prompt-based image editing that edits a source image according to a target prompt.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Kejie Wang , Xuemeng Song , Meng Liu , Jin Yuan , Weili Guan

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally limited. Training-free methods often suffer from signal…

Sound · Computer Science 2026-01-21 Ye Tao , Wen Wu , Chao Zhang , Mengyue Wu , Shuai Wang , Xuenan Xu

The increasing relevance of panoptic segmentation is tied to the advancements in autonomous driving and AR/VR applications. However, the deployment of such models has been limited due to the expensive nature of dense data annotation, giving…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Elham Amin Mansour , Ozan Unal , Suman Saha , Benjamin Bejar , Luc Van Gool

Video editing methods based on diffusion models that rely solely on a text prompt for the edit are hindered by the limited expressive power of text prompts. Thus, incorporating a reference target image as a visual guide becomes desirable…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Sai Sree Harsha , Ambareesh Revanur , Dhwanit Agarwal , Shradha Agrawal

Given an object mask, Semi-supervised Video Object Segmentation (SVOS) technique aims to track and segment the object across video frames, serving as a fundamental task in computer vision. Although recent memory-based methods demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Guanyi Qin , Ziyue Wang , Daiyun Shen , Haofeng Liu , Hantao Zhou , Junde Wu , Runze Hu , Yueming Jin

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale…

Multimedia · Computer Science 2025-01-03 Yuan Zhao , Rui Liu , Gaoxiang Cong