English
Related papers

Related papers: VideoAnydoor: High-fidelity Video Object Insertion…

200 papers

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

Robotics · Computer Science 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

We propose an efficient plug-and-play acceleration framework for semi-supervised video object segmentation by exploiting the temporal redundancies in videos presented by the compressed bitstream. Specifically, we propose a motion…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Kai Xu , Angela Yao

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge this gap often suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yujie Wei , Xinyu Liu , Shiwei Zhang , Hangjie Yuan , Jinbo Xing , Zhekai Chen , Xiang Wang , Haonan Qiu , Rui Zhao , Yutong Feng , Ruihang Chu , Yingya Zhang , Yike Guo , Xihui Liu , Hongming Shan

In this paper, we introduce a new problem of manipulating a given video by inserting other videos into it. Our main task is, given an object video and a scene video, to insert the object video at a user-specified location in the scene video…

Computer Vision and Pattern Recognition · Computer Science 2019-03-18 Donghoon Lee , Tomas Pfister , Ming-Hsuan Yang

Recent advances in diffusion-based text-to-image models have simplified creating high-fidelity images, but preserving the identity (ID) of specific elements, like a personal dog, is still challenging. Object customization, using reference…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Lingjie Kong , Kai Wu , Xiaobin Hu , Wenhui Han , Jinlong Peng , Chengming Xu , Donghao Luo , Mengtian Li , Jiangning Zhang , Chengjie Wang , Yanwei Fu

Large text-to-image diffusion models have achieved remarkable success in generating diverse, high-quality images. Additionally, these models have been successfully leveraged to edit input images by just changing the text prompt. But when…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Anant Khandelwal

Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Vladimir Kulikov , Roni Paiss , Andrey Voynov , Inbar Mosseri , Tali Dekel , Tomer Michaeli

Video depth estimation lifts monocular video clips to 3D by inferring dense depth at every frame. Recent advances in single-image depth estimation, brought about by the rise of large foundation models and the use of synthetic training data,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Bingxin Ke , Dominik Narnhofer , Shengyu Huang , Lei Ke , Torben Peters , Katerina Fragkiadaki , Anton Obukhov , Konrad Schindler

In this paper, we introduce Object-WIPER, a training-free framework for removing dynamic objects and their associated visual effects from videos, and inpainting them with semantically consistent and temporally coherent content. Our approach…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Saksham Singh Kushwaha , Sayan Nag , Yapeng Tian , Kuldeep Kulkarni

Text-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face certain limitations when dealing with personalized image…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Yuhong Zhang , Han Wang , Yiwen Wang , Rong Xie , Li Song

Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D scenes. Replicating such capability, termed compositional 3D reconstruction, is pivotal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Mingyu Dong , Chong Xia , Mingyuan Jia , Weichen Lyu , Long Xu , Zheng Zhu , Yueqi Duan

Recently, video generation has achieved significant rapid development based on superior text-to-image generation techniques. In this work, we propose a high fidelity framework for image-to-video generation, named AtomoVideo. Based on…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Litong Gong , Yiran Zhu , Weijie Li , Xiaoyang Kang , Biao Wang , Tiezheng Ge , Bo Zheng

Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object motion is essential for video generation. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zhouxia Wang , Ziyang Yuan , Xintao Wang , Tianshui Chen , Menghan Xia , Ping Luo , Ying Shan

Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Guangzhao Li , Yanming Yang , Chenxi Song , Chi Zhang

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Objects in videos are typically characterized by continuous smooth motion. We exploit continuous smooth motion in three ways. 1) Improved accuracy by using object motion as an additional source of supervision, which we obtain by…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Xin Liu , Fatemeh Karimi Nejadasl , Jan C. van Gemert , Olaf Booij , Silvia L. Pintea

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hila Chefer , Uriel Singer , Amit Zohar , Yuval Kirstain , Adam Polyak , Yaniv Taigman , Lior Wolf , Shelly Sheynin

We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal is to restore fine-grained details while keeping all…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Dewei Zhou , You Li , Zongxin Yang , Yi Yang

Video inpainting aims to fill spatio-temporal holes with plausible content in a video. Despite tremendous progress of deep neural networks for image inpainting, it is challenging to extend these methods to the video domain due to the…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Dahun Kim , Sanghyun Woo , Joon-Young Lee , In So Kweon

We propose a method for face de-identification that enables fully automatic video modification at high frame rates. The goal is to maximally decorrelate the identity, while having the perception (pose, illumination and expression) fixed. We…

Machine Learning · Computer Science 2019-11-20 Oran Gafni , Lior Wolf , Yaniv Taigman