English
Related papers

Related papers: DreamInsert: Zero-Shot Image-to-Video Object Inser…

200 papers

The task of realistically inserting a human from a reference image into a background scene is highly challenging, requiring the model to (1) determine the correct location and poses of the person and (2) perform high-quality personalization…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Jialu Gao , K J Joseph , Fernando De La Torre

Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yanchen Liu , Yanan Sun , Zhening Xing , Junyao Gao , Kai Chen , Wenjie Pei

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Songcen Xu , Hang Xu , Xiaodan Liang

Reference-based object composition involves integrating foreground reference image with background scene to produce harmonious fused image. This task becomes particularly challenging in cross-domain scenarios, where models must balance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Raghu Vamsi Chittersu , Yuvraj Singh Rathore , Pranav Adlinge , Kunal Swami

We introduce InVi, an approach for inserting or replacing objects within videos (referred to as inpainting) using off-the-shelf, text-to-image latent diffusion models. InVi targets controlled manipulation of objects and blending them…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Nirat Saini , Navaneeth Bodla , Ashish Shrivastava , Avinash Ravichandran , Xiao Zhang , Abhinav Shrivastava , Bharat Singh

We introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Ruoshi Liu , Rundi Wu , Basile Van Hoorick , Pavel Tokmakov , Sergey Zakharov , Carl Vondrick

As large-scale text-to-image generation models have made remarkable progress in the field of text-to-image generation, many fine-tuning methods have been proposed. However, these models often struggle with novel objects, especially with…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Jianxiang Lu , Cong Xie , Hui Guo

Realistic video simulation has shown significant potential across diverse applications, from virtual reality to film production. This is particularly true for scenarios where capturing videos in real-world settings is either impractical or…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Chen Bai , Zeman Shao , Guoxiang Zhang , Di Liang , Jie Yang , Zhuorui Zhang , Yujian Guo , Chengzhang Zhong , Yiqiao Qiu , Zhendong Wang , Yichen Guan , Xiaoyin Zheng , Tao Wang , Cheng Lu

Video object insertion is a critical task for dynamically inserting new objects into existing environments. Previous video generation methods focus primarily on synthesizing entire scenes while struggling with ensuring consistent object…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Xia Qi , Peishan Cong , Yichen Yao , Ziyi Wang , Yaoqin Ye , Yuexin Ma

Robotic insertion is a highly challenging task that requires exceptional precision in cluttered environments. Existing methods often have poor generalization capabilities. They typically function in restricted and structured environments,…

Robotics · Computer Science 2026-03-10 Guanghe Li , Junming Zhao , Shengjie Wang , Yang Gao

Generative video modeling has emerged as a compelling tool to zero-shot reason about plausible physical interactions for open-world manipulation. Yet, it remains a challenge to translate such human-led motions into the low-level actions…

Robotics · Computer Science 2026-01-01 Karthik Dharmarajan , Wenlong Huang , Jiajun Wu , Li Fei-Fei , Ruohan Zhang

Recent advances in customized video generation have enabled users to create videos tailored to both specific subjects and motion trajectories. However, existing methods often require complicated test-time fine-tuning and struggle with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Yujie Wei , Shiwei Zhang , Hangjie Yuan , Xiang Wang , Haonan Qiu , Rui Zhao , Yutong Feng , Feng Liu , Zhizhong Huang , Jiaxin Ye , Yingya Zhang , Hongming Shan

Novel view synthesis has observed tremendous developments since the arrival of NeRFs. However, Nerf models overfit on a single scene, lacking generalization to out of distribution objects. Recently, diffusion models have exhibited…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Rukhshanda Hussain , Hui Xian Grace Lim , Borchun Chen , Mubarak Shah , Ser Nam Lim

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Shuai Yang , Yifan Zhou , Ziwei Liu , Chen Change Loy

Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yoad Tewel , Rinon Gal , Dvir Samuel , Yuval Atzmon , Lior Wolf , Gal Chechik

Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. While diffusion models are achieving compelling results in image generation, video diffusion models are limited by…

Machine Learning · Computer Science 2024-10-28 Luca Savant Aira , Antonio Montanaro , Emanuele Aiello , Diego Valsesia , Enrico Magli

Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts fail to capture cinematic vision. This barrier limits…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Pooja Guhan , Divya Kothandaraman , Geonsun Lee , Tsung-Wei Huang , Guan-Ming Su , Dinesh Manocha

We propose ZeST, a method for zero-shot material transfer to an object in the input image given a material exemplar image. ZeST leverages existing diffusion adapters to extract implicit material representation from the exemplar image. This…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Ta-Ying Cheng , Prafull Sharma , Andrew Markham , Niki Trigoni , Varun Jampani

Text-driven object insertion in 3D scenes is an emerging task that enables intuitive scene editing through natural language. However, existing 2D editing-based methods often rely on spatial priors such as 2D masks or 3D bounding boxes, and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Chenxi Li , Weijie Wang , Qiang Li , Bruno Lepri , Nicu Sebe , Weizhi Nie

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Eyal Molad , Eliahu Horwitz , Dani Valevski , Alex Rav Acha , Yossi Matias , Yael Pritch , Yaniv Leviathan , Yedid Hoshen