English
Related papers

Related papers: Vid-CamEdit: Video Camera Trajectory Editing with …

200 papers

A free-viewpoint, editable, and high-fidelity driving simulator is crucial for training and evaluating end-to-end autonomous driving systems. In this paper, we present GA-Drive, a novel simulation framework capable of generating camera…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Hao Zhang , Lue Fan , Qitai Wang , Wenbo Li , Zehuan Wu , Lewei Lu , Zhaoxiang Zhang , Hongsheng Li

We introduce the problem of perpetual view generation - long-range generation of novel views corresponding to an arbitrarily long camera trajectory given a single image. This is a challenging problem that goes far beyond the capabilities of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Andrew Liu , Richard Tucker , Varun Jampani , Ameesh Makadia , Noah Snavely , Angjoo Kanazawa

Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, and generalized…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Xiaoming Zhao , Alex Colburn , Fangchang Ma , Miguel Angel Bautista , Joshua M. Susskind , Alexander G. Schwing

Global human motion reconstruction from in-the-wild monocular videos is increasingly demanded across VR, graphics, and robotics applications, yet requires accurate mapping of human poses from camera to world coordinates-a task challenged by…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Qijun Ying , Zhongyuan Hu , Rui Zhang , Ronghui Li , Yu Lu , Zijiao Zeng

We address the task of unsupervised retargeting of human actions from one video to another. We consider the challenging setting where only a few frames of the target is available. The core of our approach is a conditional generative model…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Jessica Lee , Deva Ramanan , Rohit Girdhar

Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the state-of-the-art of vid2vid has advanced significantly,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ting-Chun Wang , Ming-Yu Liu , Andrew Tao , Guilin Liu , Jan Kautz , Bryan Catanzaro

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

We introduce Follow-Your-Creation, a novel 4D video creation framework capable of both generating and editing 4D content from a single monocular video input. By leveraging a powerful video inpainting foundation model as a generative prior,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Yue Ma , Kunyu Feng , Xinhua Zhang , Hongyu Liu , David Junhao Zhang , Jinbo Xing , Yinhan Zhang , Ayden Yang , Zeyu Wang , Qifeng Chen

Motion-preserved video editing is crucial for creators, particularly in scenarios that demand flexibility in both the structure and semantics of swapped objects. Despite its potential, this area remains underexplored. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Sandeep Mishra , Oindrila Saha , Alan C. Bovik

Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Vladimir Kulikov , Roni Paiss , Andrey Voynov , Inbar Mosseri , Tali Dekel , Tomer Michaeli

We consider the problem of next frame prediction from video input. A recurrent convolutional neural network is trained to predict depth from monocular video input, which, along with the current video image and the camera trajectory, can…

Machine Learning · Computer Science 2017-06-14 Reza Mahjourian , Martin Wicke , Anelia Angelova

We introduce a novel camera model for monocular 3D Morphable Model (3DMM) regression methods that effectively captures the perspective distortion effect commonly seen in close-up facial images. Fitting 3D morphable models to video is a key…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Toby Chong , Ryota Nakajima

Diffusion-based approaches have recently demonstrated strong performance for single-image novel view synthesis by conditioning generative models on geometry inferred from monocular depth estimation. However, in practice, the quality and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Amirhosein Javadi , Chi-Shiang Gau , Konstantinos D. Polyzos , Tara Javidi

Recent advances in generative adversarial networks (GANs) have demonstrated the capabilities of generating stunning photo-realistic portrait images. While some prior works have applied such image GANs to unconditional 2D portrait video…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Zhongcong Xu , Jianfeng Zhang , Jun Hao Liew , Wenqing Zhang , Song Bai , Jiashi Feng , Mike Zheng Shou

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such problems typically train transformation networks to generate…

Computer Vision and Pattern Recognition · Computer Science 2018-07-27 Long Zhao , Xi Peng , Yu Tian , Mubbasir Kapadia , Dimitris Metaxas

Gaze redirection is the task of changing the gaze to a desired direction for a given monocular eye patch image. Many applications such as videoconferencing, films, games, and generation of training data for gaze estimation require…

Computer Vision and Pattern Recognition · Computer Science 2019-11-21 Zhe He , Adrian Spurr , Xucong Zhang , Otmar Hilliges

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Tengjiao Yin , Jinglei Shi , Heng Guo , Xi Wang

Generating videos for visual storytelling can be a tedious and complex process that typically requires either live-action filming or graphics animation rendering. To bypass these challenges, our key idea is to utilize the abundance of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Yingqing He , Menghan Xia , Haoxin Chen , Xiaodong Cun , Yuan Gong , Jinbo Xing , Yong Zhang , Xintao Wang , Chao Weng , Ying Shan , Qifeng Chen

A recent frontier in computer vision has been the task of 3D video generation, which consists of generating a time-varying 3D representation of a scene. To generate dynamic 3D scenes, current methods explicitly model 3D temporal dynamics by…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Rishab Parthasarathy , Zachary Ankner , Aaron Gokaslan

We focus on the task of estimating a physically plausible articulated human motion from monocular video. Existing approaches that do not consider physics often produce temporally inconsistent output with motion artifacts, while…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Erik Gärtner , Mykhaylo Andriluka , Hongyi Xu , Cristian Sminchisescu