English
Related papers

Related papers: From-Ground-To-Objects: Coarse-to-Fine Self-superv…

200 papers

Accurately perceiving location and scene is crucial for autonomous driving and mobile robots. Recent advances in deep learning have made it possible to learn egomotion and depth from monocular images in a self-supervised manner, without…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Hao Qu , Lilian Zhang , Xiaoping Hu , Xiaofeng He , Xianfei Pan , Changhao Chen

Self-supervised learning of depth map prediction and motion estimation from monocular video sequences is of vital importance -- since it realizes a broad range of tasks in robotics and autonomous vehicles. A large number of research efforts…

Computer Vision and Pattern Recognition · Computer Science 2021-03-24 Ue-Hwan Kim , Jong-Hwan Kim

Self-supervised monocular depth estimation approaches either ignore independently moving objects in the scene or need a separate segmentation step to identify them. We propose MonoDepthSeg to jointly estimate depth and segment moving…

Computer Vision and Pattern Recognition · Computer Science 2021-10-22 Sadra Safadoust , Fatma Güney

Perceiving 3D objects from monocular inputs is crucial for robotic systems, given its economy compared to multi-sensor settings. It is notably difficult as a single image can not provide any clues for predicting absolute depth values.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Tai Wang , Jiangmiao Pang , Dahua Lin

We present an end-to-end joint training framework that explicitly models 6-DoF motion of multiple dynamic objects, ego-motion and depth in a monocular camera setup without supervision. Our technical contributions are three-fold. First, we…

Computer Vision and Pattern Recognition · Computer Science 2021-02-05 Seokju Lee , Sunghoon Im , Stephen Lin , In So Kweon

Depth information is essential for on-board perception in autonomous driving and driver assistance. Monocular depth estimation (MDE) is very appealing since it allows for appearance and depth being on direct pixelwise correspondence without…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Akhil Gurram , Ahmet Faruk Tuna , Fengyi Shen , Onay Urfalioglu , Antonio M. López

Monocular depth estimation has been increasingly adopted in robotics and autonomous driving for its ability to infer scene geometry from a single camera. In self-supervised monocular depth estimation frameworks, the network jointly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Tae-Wook Um , Ki-Hyeon Kim , Hyun-Duck Choi , Hyo-Sung Ahn

Monocular depth estimation (MDE) aims to transform an RGB image of a scene into a pixelwise depth map from the same camera view. It is fundamentally ill-posed due to missing information: any single image can have been taken from many…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Dylan Auty , Krystian Mikolajczyk

Monocular Depth Estimation (MDE) enables spatial understanding, 3D reconstruction, and autonomous navigation, yet deep learning approaches often predict only relative depth without a consistent metric scale. This limitation reduces…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Jiuling Zhang

Estimating the 3D position and orientation of objects in the environment with a single RGB camera is a critical and challenging task for low-cost urban autonomous driving and mobile robots. Most of the existing algorithms are based on the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Yuxuan Liu , Yuan Yixuan , Ming Liu

We present a self-supervised learning framework to estimate the individual object motion and monocular depth from video. We model the object motion as a 6 degree-of-freedom rigid-body transformation. The instance segmentation mask is…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Qi Dai , Vaishakh Patil , Simon Hecker , Dengxin Dai , Luc Van Gool , Konrad Schindler

Monocular Depth Estimation (MDE) plays a vital role in applications such as autonomous driving. However, various attacks target MDE models, with physical attacks posing significant threats to system security. Traditional adversarial…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Zhiyuan Cheng , Cheng Han , James Liang , Qifan Wang , Xiangyu Zhang , Dongfang Liu

Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-relevant target regions from the surrounding context. We…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yuxin Du , Tao Lin , Zile Zhong , Runting Li , Xiyao Chen , Jiting Liu , Chenglin Liu , Ying-Cong Chen , Yuqian Fu , Bo Zhao

Recently, self-supervised monocular depth estimation has gained popularity with numerous applications in autonomous driving and robotics. However, existing solutions primarily seek to estimate depth from immediate visual features, and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Youhong Wang , Yunji Liang , Hao Xu , Shaohui Jiao , Hongkai Yu

As an agent moves through the world, the apparent motion of scene elements is (usually) inversely proportional to their depth. It is natural for a learning agent to associate image patterns with the magnitude of their displacement over…

Computer Vision and Pattern Recognition · Computer Science 2018-04-03 Huaizu Jiang , Erik Learned-Miller , Gustav Larsson , Michael Maire , Greg Shakhnarovich

Recent camera-based 3D object detection is limited by the precision of transforming from image to 3D feature spaces, as well as the accuracy of object localization within the 3D space. This paper aims to address such a fundamental problem…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Chaoqun Wang , Yiran Qin , Zijian Kang , Ningning Ma , Ruimao Zhang

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Jaehoon Choi , Dongki Jung , Yonghan Lee , Deokhwa Kim , Dinesh Manocha , Donghwan Lee

Existing methods for scale-invariant monocular depth estimation (SI MDE) often struggle due to the complexity of the task, and limited and non-diverse datasets, hindering generalizability in real-world scenarios. This is while…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 S. Mahdi H. Miangoleh , Mahesh Reddy , Yağız Aksoy

This paper proposes a self-supervised monocular image-to-depth prediction framework that is trained with an end-to-end photometric loss that handles not only 6-DOF camera motion but also 6-DOF moving object instances. Self-supervision is…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Houssem Boulahbal , Adrian Voicila , Andrew Comport

Monocular depth estimation (MDE) in the self-supervised scenario has emerged as a promising method as it refrains from the requirement of ground truth depth. Despite continuous efforts, MDE is still sensitive to scale changes especially…

Computer Vision and Pattern Recognition · Computer Science 2023-02-21 Peizhe Jiang , Wei Yang , Xiaoqing Ye , Xiao Tan , Meng Wu