English
Related papers

Related papers: DO3D: Self-supervised Learning of Decomposed Objec…

200 papers

Depth from a monocular video can enable billions of devices and robots with a single camera to see the world in 3D. In this paper, we present an approach with a differentiable flow-to-depth layer for video depth estimation. The model…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Jiaxin Xie , Chenyang Lei , Zhuwen Li , Li Erran Li , Qifeng Chen

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu

Accurately analyzing the motion parts and their motion attributes in dynamic environments is crucial for advancing key areas such as embodied intelligence. Addressing the limitations of existing methods that rely on dense multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Hongyi Zhou , Yulan Guo , Xiaogang Wang , Kai Xu

Self-supervised monocular depth estimation is a significant task for low-cost and efficient 3D scene perception and measurement in endoscopy. However, the variety of illumination conditions and scene features is still the primary challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Liangjing Shao , Chenkang Du , Benshuang Chen , Xueli Liu , Xinrong Chen

We present a method for decomposing the 3D scene flow observed from a moving stereo rig into stationary scene elements and dynamic object motion. Our unsupervised learning framework jointly reasons about the camera motion, optical flow, and…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Seokju Lee , Sunghoon Im , Stephen Lin , In So Kweon

Self-supervised depth learning from monocular images normally relies on the 2D pixel-wise photometric relation between temporally adjacent image frames. However, they neither fully exploit the 3D point-wise geometric correspondences, nor…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Kaichen Zhou , Lanqing Hong , Changhao Chen , Hang Xu , Chaoqiang Ye , Qingyong Hu , Zhenguo Li

In this paper we present a novel self-supervised method to anticipate the depth estimate for a future, unobserved real-world urban scene. This work is the first to explore self-supervised learning for estimation of monocular depth of future…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Sauradip Nag , Nisarg Shah , Anran Qi , Raghavendra Ramachandra

Monocular 3D detection has drawn much attention from the community due to its low cost and setup simplicity. It takes an RGB image as input and predicts 3D boxes in the 3D space. The most challenging sub-task lies in the instance depth…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Liang Peng , Xiaopei Wu , Zheng Yang , Haifeng Liu , Deng Cai

Monocular 3D lane detection is essential for autonomous driving, but challenging due to the inherent lack of explicit spatial information. Multi-modal approaches rely on expensive depth sensors, while methods incorporating fully-supervised…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Max van den Hoven , Kishaan Jeeveswaran , Pieter Piscaer , Thijs Wensveen , Elahe Arani , Bahram Zonooz

Monocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image. In this paper, we strive to boost currently underperforming monocular…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Weijia Zhang , Dongnan Liu , Chao Ma , Weidong Cai

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hao Zhang , Fang Li , Samyak Rawlekar , Narendra Ahuja

We present a self-supervised learning-based pipeline for dense 3D reconstruction from full-length monocular endoscopic videos without a priori modeling of anatomy or shading. Our method only relies on unlabeled monocular endoscopic videos…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Xingtong Liu , Ayushi Sinha , Masaru Ishii , Gregory D. Hager , Russell H. Taylor , Mathias Unberath

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth…

Robotics · Computer Science 2022-11-29 Ruofeng Wei , Bin Li , Hangjie Mo , Fangxun Zhong , Yonghao Long , Qi Dou , Yun-Hui Liu , Dong Sun

We present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or shading. Our method only requires monocular endoscopic videos…

Computer Vision and Pattern Recognition · Computer Science 2019-10-31 Xingtong Liu , Ayushi Sinha , Masaru Ishii , Gregory D. Hager , Austin Reiter , Russell H. Taylor , Mathias Unberath

In this paper, we tackle the problem of estimating the depth of a scene from a monocular video sequence. In particular, we handle challenging scenarios, such as non-translational camera motion and dynamic scenes, where traditional structure…

Computer Vision and Pattern Recognition · Computer Science 2015-11-20 Miaomiao Liu , Mathieu Salzmann , Xuming He

A thermal camera can robustly capture thermal radiation images under harsh light conditions such as night scenes, tunnels, and disaster scenarios. However, despite this advantage, neither depth nor ego-motion estimation research for the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Ukcheol Shin , Kyunghyun Lee , Seokju Lee , In So Kweon

Self-supervised monocular depth estimation (MDE) has gained popularity for obtaining depth predictions directly from videos. However, these methods often produce scale invariant results, unless additional training signals are provided.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Gasser Elazab , Torben Gräber , Michael Unterreiner , Olaf Hellwich

Self-supervised monocular depth estimation is a salient task for 3D scene understanding. Learned jointly with monocular ego-motion estimation, several methods have been proposed to predict accurate pixel-wise depth without using labeled…

Computer Vision and Pattern Recognition · Computer Science 2023-02-02 Hemang Chawla , Kishaan Jeeveswaran , Elahe Arani , Bahram Zonooz

Supervised learning based methods for monocular depth estimation usually require large amounts of extensively annotated training data. In the case of aerial imagery, this ground truth is particularly difficult to acquire. Therefore, in this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Max Hermann , Boitumelo Ruf , Martin Weinmann , Stefan Hinz

Recently, self-supervised learning technology has been applied to calculate depth and ego-motion from monocular videos, achieving remarkable performance in autonomous driving scenarios. One widely adopted assumption of depth and ego-motion…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Shuwei Shao , Zhongcai Pei , Weihai Chen , Wentao Zhu , Xingming Wu , Dianmin Sun , Baochang Zhang