English
Related papers

Related papers: Place Anything into Any Video

200 papers

We propose a novel unsupervised method to autoregressively generate videos from a single frame and a sparse motion input. Our trained model can generate unseen realistic object-to-object interactions. Although our model has never been given…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Aram Davtyan , Paolo Favaro

Generating dynamic 3D object from a single-view video is challenging due to the lack of 4D labeled data. An intuitive approach is to extend previous image-to-3D pipelines by transferring off-the-shelf image generation models such as score…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Zijie Pan , Zeyu Yang , Xiatian Zhu , Li Zhang

We propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary videos. While seemingly easy for humans, this problem poses many challenges for computers. We propose to approach this problem by…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Shengyi Qian , Linyi Jin , Chris Rockwell , Siyi Chen , David F. Fouhey

Existing person video generation methods either lack the flexibility in controlling both the appearance and motion, or fail to preserve detailed appearance and temporal consistency. In this paper, we tackle the problem of motion transfer…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Kun Cheng , Hao-Zhi Huang , Chun Yuan , Lingyiqing Zhou , Wei Liu

To watch 360{\deg} videos on normal 2D displays, we need to project the selected part of the 360{\deg} image onto the 2D display plane. In this paper, we propose a fully-automated framework for generating content-aware 2D normal-view…

Graphics · Computer Science 2017-09-12 Yeong Won Kim , Dae-Yong Jo , Chang-Ryeol Lee , Hyeok-Jae Choi , Yong Hoon Kwon , Kuk-Jin Yoon

This work presents AnyDoor, a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations in a harmonious way. Instead of tuning parameters for each object, our model is trained only…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Xi Chen , Lianghua Huang , Yu Liu , Yujun Shen , Deli Zhao , Hengshuang Zhao

We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method aims to predict open-ended human action changes in video.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Lan Wang , Vishnu Boddeti , Sernam Lim

Image generation and editing have seen a great deal of advancements with the rise of large-scale diffusion models that allow user control of different modalities such as text, mask, depth maps, etc. However, controlled editing of videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 AmirHossein Zamani , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

One key challenge in augmented reality is the placement of virtual content in natural locations. Existing automated techniques are only able to work with a closed-vocabulary, fixed set of objects. In this paper, we introduce a new…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Luke Yoffe , Aditya Sharma , Tobias Höllerer

While the satellite-based Global Positioning System (GPS) is adequate for some outdoor applications, many other applications are held back by its multi-meter positioning errors and poor indoor coverage. In this paper, we study the…

Computer Vision and Pattern Recognition · Computer Science 2020-02-20 Abm Musa , Jakob Eriksson

In this work, we present a novel approach for motion customization in video generation, addressing the widespread gap in the exploration of motion representation within video generative models. Recognizing the unique challenges posed by the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Luozhou Wang , Ziyang Mai , Guibao Shen , Yixun Liang , Xin Tao , Pengfei Wan , Di Zhang , Yijun Li , Yingcong Chen

Given a demonstration of a complex manipulation task, such as pouring liquid from one container to another, we seek to generate a motion plan for a new task instance involving objects with different geometries. This is nontrivial since we…

We introduce PhotoDoodle, a novel image editing framework designed to facilitate photo doodling by enabling artists to overlay decorative elements onto photographs. Photo doodling is challenging because the inserted elements must appear…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Shijie Huang , Yiren Song , Yuxuan Zhang , Hailong Guo , Xueyin Wang , Mike Zheng Shou , Jiaming Liu

Video inpainting aims to fill spatio-temporal holes with plausible content in a video. Despite tremendous progress of deep neural networks for image inpainting, it is challenging to extend these methods to the video domain due to the…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Dahun Kim , Sanghyun Woo , Joon-Young Lee , In So Kweon

Reconstructing 3D object models is playing an important role in many applications in the field of computer vision. Instead of employing a collection of cameras and/or sensors as in many studies, this paper proposes a simple way to build a…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Trong Nguyen Nguyen , Huu Hung Huynh , Jean Meunier

The majority of modern robot learning methods focus on learning a set of pre-defined tasks with limited or no generalization to new tasks. Extending the robot skillset to novel tasks involves gathering an extensive amount of training data…

Robotics · Computer Science 2025-04-03 Dandan Shan , Kaichun Mo , Wei Yang , Yu-Wei Chao , David Fouhey , Dieter Fox , Arsalan Mousavian

While image manipulation achieves tremendous breakthroughs (e.g., generating realistic faces) in recent years, video generation is much less explored and harder to control, which limits its applications in the real world. For instance,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-08 Tsun-Hsuan Wang , Yen-Chi Cheng , Chieh Hubert Lin , Hwann-Tzong Chen , Min Sun

Monocular depth estimation is crucial for tracking and reconstruction algorithms, particularly in the context of surgical videos. However, the inherent challenges in directly obtaining ground truth depth maps during surgery render…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ange Lou , Yamin Li , Yike Zhang , Jack Noble

We present Image Sculpting, a new framework for editing 2D images by incorporating tools from 3D geometry and graphics. This approach differs markedly from existing methods, which are confined to 2D spaces and typically rely on textual…

Graphics · Computer Science 2024-01-04 Jiraphon Yenphraphai , Xichen Pan , Sainan Liu , Daniele Panozzo , Saining Xie

Humans use context and scene knowledge to easily localize moving objects in conditions of complex illumination changes, scene clutter and occlusions. In this paper, we present a method to leverage human knowledge in the form of annotated…

Computer Vision and Pattern Recognition · Computer Science 2016-04-20 Archith J. Bency , S. Karthikeyan , Carter De Leo , Santhoshkumar Sunderrajan , B. S. Manjunath
‹ Prev 1 8 9 10 Next ›