English
Related papers

Related papers: Unsupervised Video Prediction from a Single Frame …

200 papers

Video prediction, forecasting the future frames from a sequence of input frames, is a challenging task since the view changes are influenced by various factors, such as the global context surrounding the scene and local motion dynamics. In…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Jaehoon Cho , Jiyoung Lee , Changjae Oh , Wonil Song , Kwanghoon Sohn

We present an unsupervised learning framework for the task of monocular depth and camera motion estimation from unstructured video sequences. We achieve this by simultaneously training depth and camera pose estimation networks using the…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Tinghui Zhou , Matthew Brown , Noah Snavely , David G. Lowe

We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Rong Wang , Ruyi Zha , Ziang Cheng , Jiayu Yang , Pulak Purkait , Hongdong Li

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jingxi Chen , Zongxia Li , Zhichao Liu , Guangyao Shi , Xiyang Wu , Fuxiao Liu , Cornelia Fermuller , Brandon Y. Feng , Yiannis Aloimonos

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

We introduce VividDream, a method for generating explorable 4D scenes with ambient dynamics from a single input image or text prompt. VividDream first expands an input image into a static 3D point cloud through iterative inpainting and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Yao-Chih Lee , Yi-Ting Chen , Andrew Wang , Ting-Hsuan Liao , Brandon Y. Feng , Jia-Bin Huang

We propose a self-supervised visual learning method by predicting the variable playback speeds of a video. Without semantic labels, we learn the spatio-temporal visual representation of the video by leveraging the variations in the visual…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Hyeon Cho , Taehoon Kim , Hyung Jin Chang , Wonjun Hwang

Large models have shown generalization across datasets for many low-level vision tasks, like depth estimation, but no such general models exist for scene flow. Even though scene flow has wide potential use, it is not used in practice…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Yiqing Liang , Abhishek Badki , Hang Su , James Tompkin , Orazio Gallo

In this paper, we introduce a novel framework that can learn to make visual predictions about the motion of a robotic agent from raw video frames. Our proposed motion prediction network (PROM-Net) can learn in a completely unsupervised…

Robotics · Computer Science 2019-06-26 Meenakshi Sarkar , Prabhu Pradhan , Debasish Ghose

Envisioning physically plausible outcomes from a single image requires a deep understanding of the world's dynamics. To address this, we introduce PhysGen3D, a novel framework that transforms a single image into an amodal, camera-centric,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Boyuan Chen , Hanxiao Jiang , Shaowei Liu , Saurabh Gupta , Yunzhu Li , Hao Zhao , Shenlong Wang

Unsupervised learning of a generalizable model of the visual appearance of humans from video data is of major importance for computing systems interacting naturally with their users and others. We propose a step towards automatic behavior…

Computer Vision and Pattern Recognition · Computer Science 2017-04-13 Thomas Walther , Rolf P. Würtz

We present a learning-based model to infer the personalized 3D shape of people from a few frames (1-8) of a monocular video in which the person is moving, in less than 10 seconds with a reconstruction accuracy of 5mm. Our model learns to…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Thiemo Alldieck , Marcus Magnor , Bharat Lal Bhatnagar , Christian Theobalt , Gerard Pons-Moll

We propose an efficient approach to exploiting motion information from consecutive frames of a video sequence to recover the 3D pose of people. Previous approaches typically compute candidate poses in individual frames and then link them in…

Computer Vision and Pattern Recognition · Computer Science 2016-09-05 Bugra Tekin , Artem Rozantsev , Vincent Lepetit , Pascal Fua

Estimating 3D scene flow from a sequence of monocular images has been gaining increased attention due to the simple, economical capture setup. Owing to the severe ill-posedness of the problem, the accuracy of current methods has been…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Junhwa Hur , Stefan Roth

Existing deep models predict 2D and 3D kinematic poses from video that are approximately accurate, but contain visible errors that violate physical constraints, such as feet penetrating the ground and bodies leaning at extreme angles. In…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Davis Rempe , Leonidas J. Guibas , Aaron Hertzmann , Bryan Russell , Ruben Villegas , Jimei Yang

Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework of…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Songlin Wei , Guodong Chen , Wenzheng Chi , Zhenhua Wang , Lining Sun

This paper presents a new self-supervised system for learning to detect novel and previously unseen categories of objects in images. The proposed system receives as input several unlabeled videos of scenes containing various objects. The…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Juntao Tan , Changkyu Song , Abdeslam Boularias

Successful video analysis relies on accurate recognition of pixels across frames, and frame reconstruction methods based on video correspondence learning are popular due to their efficiency. Existing frame reconstruction methods, while…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Zihan Zhou , Changrui Dai , Aibo Song , Xiaolin Fang

We propose a data-driven scene flow estimation algorithm exploiting the observation that many 3D scenes can be explained by a collection of agents moving as rigid bodies. At the core of our method lies a deep architecture able to reason at…

Computer Vision and Pattern Recognition · Computer Science 2021-02-18 Zan Gojcic , Or Litany , Andreas Wieser , Leonidas J. Guibas , Tolga Birdal

Future frame prediction in videos is a challenging problem because videos include complicated movements and large appearance changes. Learning-based future frame prediction approaches have been proposed in kinds of literature. A common…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Wonjik Kim , Masayuki Tanaka , Masatoshi Okutomi , Yoko Sasaki
‹ Prev 1 8 9 10 Next ›