English
Related papers

Related papers: Towards Scale-Aware, Robust, and Generalizable Uns…

200 papers

Self-supervised monocular depth estimation networks are trained to predict scene depth using nearby frames as a supervision signal during training. However, for many applications, sequence information in the form of video frames is also…

Computer Vision and Pattern Recognition · Computer Science 2021-07-15 Jamie Watson , Oisin Mac Aodha , Victor Prisacariu , Gabriel Brostow , Michael Firman

With the development of computational intelligence algorithms, unsupervised monocular depth and pose estimation framework, which is driven by warped photometric consistency, has shown great performance in the daytime scenario. While in some…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Chaoqiang Zhao , Yang Tang , Qiyu Sun

Self-supervised monocular depth estimation has gathered notable interest since it can liberate training from dependency on depth annotations. In monocular video training case, recent methods only conduct view synthesis between existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Jinfeng Liu , Lingtong Kong , Bo Li , Zerong Wang , Hong Gu , Jinwei Chen

Monocular metric depth estimation (MMDE) is a core challenge in computer vision, playing a pivotal role in real-world applications that demand accurate spatial understanding. Although prior works have shown promising zero-shot performance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Girish Chandar Ganesan , Yuliang Guo , Liu Ren , Xiaoming Liu

Recent advances in self-supervised learning havedemonstrated that it is possible to learn accurate monoculardepth reconstruction from raw video data, without using any 3Dground truth for supervision. However, in robotics…

Computer Vision and Pattern Recognition · Computer Science 2020-04-14 Robert McCraith , Lukas Neumann , Andrew Zisserman , Andrea Vedaldi

An event camera is a novel vision sensor that can capture per-pixel brightness changes and output a stream of asynchronous ``events''. It has advantages over conventional cameras in those scenes with high-speed motions and challenging…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Junyu Zhu , Lina Liu , Bofeng Jiang , Feng Wen , Hongbo Zhang , Wanlong Li , Yong Liu

Aerial scene understanding systems face stringent payload restrictions and must often rely on monocular depth estimation for modeling scene geometry, which is an inherently ill-posed problem. Moreover, obtaining accurate ground truth data…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Horatiu Florea , Sergiu Nedevschi

Self-supervised monocular depth estimation (MDE) has gained popularity for obtaining depth predictions directly from videos. However, these methods often produce scale invariant results, unless additional training signals are provided.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Gasser Elazab , Torben Gräber , Michael Unterreiner , Olaf Hellwich

Self-supervised methods have showed promising results on depth estimation task. However, previous methods estimate the target depth map and camera ego-motion simultaneously, underusing multi-frame correlation information and ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Songchun Zhang , Chunhui Zhao

Self-supervised monocular scene flow estimation, aiming to understand both 3D structures and 3D motions from two temporally consecutive monocular images, has received increasing attention for its simple and economical sensor setup. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Zijie Jiang , Masatoshi Okutomi

We propose a novel monocular visual odometry (VO) system called UnDeepVO in this paper. UnDeepVO is able to estimate the 6-DoF pose of a monocular camera and the depth of its view by using deep neural networks. There are two salient…

Computer Vision and Pattern Recognition · Computer Science 2018-02-22 Ruihao Li , Sen Wang , Zhiqiang Long , Dongbing Gu

We present a novel method for simultaneous learning of depth, egomotion, object motion, and camera intrinsics from monocular videos, using only consistency across neighboring video frames as supervision signal. Similarly to prior work, our…

Computer Vision and Pattern Recognition · Computer Science 2019-10-31 Ariel Gordon , Hanhan Li , Rico Jonschkowski , Anelia Angelova

The recent development of \emph{foundation models} for monocular depth estimation such as Depth Anything paved the way to zero-shot monocular depth estimation. Since it returns an affine-invariant disparity map, the favored technique to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Rémi Marsal , Alexandre Chapoutot , Philippe Xu , David Filliat

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Jaehoon Choi , Dongki Jung , Yonghan Lee , Deokhwa Kim , Dinesh Manocha , Donghwan Lee

We present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion and depth in a monocular camera setup without supervision. Our technical contributions are three-fold. First, we…

Computer Vision and Pattern Recognition · Computer Science 2020-04-09 Seokju Lee , Sunghoon Im , Stephen Lin , In So Kweon

Conventional self-supervised monocular depth prediction methods are based on a static environment assumption, which leads to accuracy degradation in dynamic scenes due to the mismatch and occlusion problems introduced by object motions.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Ziyue Feng , Liang Yang , Longlong Jing , Haiyan Wang , YingLi Tian , Bing Li

As a flexible passive 3D sensing means, unsupervised learning of depth from monocular videos is becoming an important research topic. It utilizes the photometric errors between the target view and the synthesized views from its adjacent…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Hualie Jiang , Laiyan Ding , Zhenglong Sun , Rui Huang

Self-supervised multi-frame monocular depth estimation relies on the geometric consistency between successive frames under the assumption of a static scene. However, the presence of moving objects in dynamic scenes introduces inevitable…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Sungmin Woo , Wonjoon Lee , Woo Jin Kim , Dogyoon Lee , Sangyoun Lee

This paper presents an self-supervised deep learning network for monocular visual inertial odometry (named DeepVIO). DeepVIO provides absolute trajectory estimation by directly merging 2D optical flow feature (OFF) and Inertial Measurement…

Robotics · Computer Science 2019-07-01 Liming Han , Yimin Lin , Guoguang Du , Shiguo Lian

Fisheye cameras are commonly used in applications like autonomous driving and surveillance to provide a large field of view ($>180^{\circ}$). However, they come at the cost of strong non-linear distortions which require more complex…

Computer Vision and Pattern Recognition · Computer Science 2020-10-08 Varun Ravi Kumar , Sandesh Athni Hiremath , Stefan Milz , Christian Witt , Clement Pinnard , Senthil Yogamani , Patrick Mader