English
Related papers

Related papers: Epipolar Geometry based Learning of Multi-view Dep…

200 papers

Scale-aware monocular depth estimation poses a significant challenge in computer-aided endoscopic navigation. However, existing depth estimation methods that do not consider the geometric priors struggle to learn the absolute scale from…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Ruofeng Wei , Bin Li , Kai Chen , Yiyao Ma , Yunhui Liu , Qi Dou

We present a novel method to train machine learning algorithms to estimate scene depths from a single image, by using the information provided by a camera's aperture as supervision. Prior works use a depth sensor's outputs or images of the…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Pratul P. Srinivasan , Rahul Garg , Neal Wadhwa , Ren Ng , Jonathan T. Barron

We present an unsupervised learning framework for the task of monocular depth and camera motion estimation from unstructured video sequences. We achieve this by simultaneously training depth and camera pose estimation networks using the…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Tinghui Zhou , Matthew Brown , Noah Snavely , David G. Lowe

Self-supervised monocular depth estimation (MDE) has gained popularity for obtaining depth predictions directly from videos. However, these methods often produce scale invariant results, unless additional training signals are provided.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Gasser Elazab , Torben Gräber , Michael Unterreiner , Olaf Hellwich

Accurate relative pose is one of the key components in visual odometry (VO) and simultaneous localization and mapping (SLAM). Recently, the self-supervised learning framework that jointly optimizes the relative pose and target image depth…

Computer Vision and Pattern Recognition · Computer Science 2019-02-26 Tianwei Shen , Zixin Luo , Lei Zhou , Hanyu Deng , Runze Zhang , Tian Fang , Long Quan

In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance are based on geometry and have to be carefully designed for…

Computer Vision and Pattern Recognition · Computer Science 2020-02-19 Huangying Zhan , Chamara Saroj Weerasekera , Jiawang Bian , Ian Reid

We propose DFPNet -- an unsupervised, joint learning system for monocular Depth, Optical Flow and egomotion (Camera Pose) estimation from monocular image sequences. Due to the nature of 3D scene geometry these three components are coupled.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Dipan Mandal , Abhilash Jain

Monocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in terms of evaluation metrics such as the pixel-wise relative error, most methods neglect the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Wei Yin , Yifan Liu , Chunhua Shen

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

In this paper we formulate structure from motion as a learning problem. We train a convolutional network end-to-end to compute depth and camera motion from successive, unconstrained image pairs. The architecture is composed of multiple…

Computer Vision and Pattern Recognition · Computer Science 2018-01-18 Benjamin Ummenhofer , Huizhong Zhou , Jonas Uhrig , Nikolaus Mayer , Eddy Ilg , Alexey Dosovitskiy , Thomas Brox

Depth estimation in surgical video plays a crucial role in many image-guided surgery procedures. However, it is difficult and time consuming to create depth map ground truth datasets in surgical videos due in part to inconsistent brightness…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Ange Lou , Jack Noble

We consider the problem of next frame prediction from video input. A recurrent convolutional neural network is trained to predict depth from monocular video input, which, along with the current video image and the camera trajectory, can…

Machine Learning · Computer Science 2017-06-14 Reza Mahjourian , Martin Wicke , Anelia Angelova

Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted depth is increasingly used as an input signal for downstream tasks such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Viktor Kocur , Sithu Aung , Gabrielle Flood , Yaqing Ding , Lukas Bujnak , Torsten Sattler , Zuzana Kukelova

Learning to predict scene depth and camera motion from RGB inputs only is a challenging task. Most existing learning based methods deal with this task in a supervised manner which require ground-truth data that is expensive to acquire. More…

Computer Vision and Pattern Recognition · Computer Science 2019-10-01 Yunxiao Shi , Jing Zhu , Yi Fang , Kuochin Lien , Junli Gu

A thermal camera can robustly capture thermal radiation images under harsh light conditions such as night scenes, tunnels, and disaster scenarios. However, despite this advantage, neither depth nor ego-motion estimation research for the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Ukcheol Shin , Kyunghyun Lee , Seokju Lee , In So Kweon

Supervised learning based methods for monocular depth estimation usually require large amounts of extensively annotated training data. In the case of aerial imagery, this ground truth is particularly difficult to acquire. Therefore, in this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Max Hermann , Boitumelo Ruf , Martin Weinmann , Stefan Hinz

Omnidirectional 3D information is essential for a wide range of applications such as Virtual Reality, Autonomous Driving, Robotics, etc. In this paper, we propose a novel, model-agnostic, two-stage pipeline for omnidirectional monocular…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Yuyan Li , Zhixin Yan , Ye Duan , Liu Ren

Self-supervised monocular depth estimation has been widely investigated to estimate depth images and relative poses from RGB images. This framework is attractive for researchers because the depth and pose networks can be trained from just…

Computer Vision and Pattern Recognition · Computer Science 2022-02-21 Noriaki Hirose , Kosuke Tahara

Video generation models have progressed tremendously through large latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Orest Kupyn , Fabian Manhardt , Federico Tombari , Christian Rupprecht

Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Clément Godard , Oisin Mac Aodha , Michael Firman , Gabriel Brostow