English
Related papers

Related papers: Vision-Language Embodiment for Monocular Depth Est…

200 papers

Spatial visual perception is a fundamental requirement in physical-world applications like autonomous driving and robotic manipulation, driven by the need to interact with 3D environments. Capturing pixel-aligned metric depth using RGB-D…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Bin Tan , Changjiang Sun , Xiage Qin , Hanat Adai , Zelin Fu , Tianxiang Zhou , Han Zhang , Yinghao Xu , Xing Zhu , Yujun Shen , Nan Xue

Spatial scene-understanding, including dense depth and ego-motion estimation, is an important problem in computer vision for autonomous vehicles and advanced driver assistance systems. Thus, it is beneficial to design perception modules…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Hemang Chawla , Matti Jukola , Shabbir Marzban , Elahe Arani , Bahram Zonooz

Monocular depth estimation is scale-ambiguous, and thus requires scale supervision to produce metric predictions. Even so, the resulting models will be geometry-specific, with learned scales that cannot be directly transferred across…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Vitor Guizilini , Igor Vasiljevic , Dian Chen , Rares Ambrus , Adrien Gaidon

Supervised learning based methods for monocular depth estimation usually require large amounts of extensively annotated training data. In the case of aerial imagery, this ground truth is particularly difficult to acquire. Therefore, in this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Max Hermann , Boitumelo Ruf , Martin Weinmann , Stefan Hinz

In the recent years, many methods demonstrated the ability of neural networks to learn depth and pose changes in a sequence of images, using only self-supervision as the training signal. Whilst the networks achieve good performance, the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 Robert McCraith , Lukas Neumann , Andrea Vedaldi

Current approaches to semantic image and scene understanding typically employ rather simple object representations such as 2D or 3D bounding boxes. While such coarse models are robust and allow for reliable object detection, they discard…

Computer Vision and Pattern Recognition · Computer Science 2014-11-24 M. Zeeshan Zia , Michael Stark , Konrad Schindler

This paper addresses the problem of estimating the depth map of a scene given a single RGB image. We propose a fully convolutional architecture, encompassing residual learning, to model the ambiguous mapping between monocular images and…

Computer Vision and Pattern Recognition · Computer Science 2016-09-20 Iro Laina , Christian Rupprecht , Vasileios Belagiannis , Federico Tombari , Nassir Navab

Accurately estimating depth in 360-degree imagery is crucial for virtual reality, autonomous navigation, and immersive media applications. Existing depth estimation methods designed for perspective-view imagery fail when applied to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Ning-Hsu Wang , Yu-Lun Liu

Self-supervised monocular depth estimation (SSMDE) has gained attention in the field of deep learning as it estimates depth without requiring ground truth depth maps. This approach typically uses a photometric consistency loss between a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Wonhyeok Choi , Kyumin Hwang , Minwoo Choi , Kiljoon Han , Wonjoon Choi , Mingyu Shin , Sunghoon Im

We present a method to estimate depth of a dynamic scene, containing arbitrary moving objects, from an ordinary video captured with a moving camera. We seek a geometrically and temporally consistent solution to this underconstrained…

Computer Vision and Pattern Recognition · Computer Science 2021-08-04 Zhoutong Zhang , Forrester Cole , Richard Tucker , William T. Freeman , Tali Dekel

Embodied scene understanding serves as the cornerstone for autonomous agents to perceive, interpret, and respond to open driving scenarios. Such understanding is typically founded upon Vision-Language Models (VLMs). Nevertheless, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Yunsong Zhou , Linyan Huang , Qingwen Bu , Jia Zeng , Tianyu Li , Hang Qiu , Hongzi Zhu , Minyi Guo , Yu Qiao , Hongyang Li

We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

We consider the problem of depth estimation from a single monocular image in this work. It is a challenging task as no reliable depth cues are available, e.g., stereo correspondences, motions, etc. Previous efforts have been focusing on…

Computer Vision and Pattern Recognition · Computer Science 2015-10-01 Fayao Liu , Chunhua Shen , Guosheng Lin

Traditional monocular depth estimation suffers from inherent ambiguity and visual nuisances. We demonstrate that language can enhance monocular depth estimation by providing an additional condition (rather than images alone) aligned with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Ziyao Zeng , Jingcheng Ni , Daniel Wang , Patrick Rim , Younjoon Chung , Fengyu Yang , Byung-Woo Hong , Alex Wong

The dense depth estimation of a 3D scene has numerous applications, mainly in robotics and surveillance. LiDAR and radar sensors are the hardware solution for real-time depth estimation, but these sensors produce sparse depth maps and are…

Computer Vision and Pattern Recognition · Computer Science 2021-03-02 Alwyn Mathew , Aditya Prakash Patra , Jimson Mathew

360{\deg} cameras can capture complete environments in a single shot, which makes 360{\deg} imagery alluring in many computer vision tasks. However, monocular depth estimation remains a challenge for 360{\deg} data, particularly for high…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Manuel Rey-Area , Mingze Yuan , Christian Richardt

A metric-accurate semantic 3D representation is essential for many robotic tasks. This work proposes a simple, yet powerful, way to integrate the 2D embeddings of a Vision-Language Model in a metric-accurate 3D representation at real-time.…

Robotics · Computer Science 2025-08-11 Christian Rauch , Björn Ellensohn , Linus Nwankwo , Vedant Dave , Elmar Rueckert

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance,…

Self-supervised monocular depth estimation holds significant importance in the fields of autonomous driving and robotics. However, existing methods are typically trained and tested on standard datasets, overlooking the impact of various…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Ziyang Song , Ruijie Zhu , Chuxin Wang , Jiacheng Deng , Jianfeng He , Tianzhu Zhang

Monocular depth estimation and defocus estimation are two fundamental tasks in computer vision. Most existing methods treat depth estimation and defocus estimation as two separate tasks, ignoring the strong connection between them. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Renzhi He , Hualin Hong , Boya Fu , Fei Liu