English
Related papers

Related papers: WorDepth: Variational Language Prior for Monocular…

200 papers

Depth information is the foundation of perception, essential for autonomous driving, robotics, and other source-constrained applications. Promptly obtaining accurate and efficient depth information allows for a rapid response in dynamic…

Computer Vision and Pattern Recognition · Computer Science 2022-10-26 Xin Zhang , Rabab Abdelfattah , Yuqi Song , Samuel A. Dauchert , Xiaofeng wang

We propose a monocular depth estimation method based on visual autoregressive (VAR) priors, offering an alternative to diffusion-based approaches. Our method adapts a large-scale text-to-image VAR model and introduces a scale-wise…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Amir El-Ghoussani , André Kaup , Nassir Navab , Gustavo Carneiro , Vasileios Belagiannis

Monocular 3D detection has drawn much attention from the community due to its low cost and setup simplicity. It takes an RGB image as input and predicts 3D boxes in the 3D space. The most challenging sub-task lies in the instance depth…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Liang Peng , Xiaopei Wu , Zheng Yang , Haifeng Liu , Deng Cai

Monocular depth estimation is the base task in computer vision. It has a tremendous development in the decade with the development of deep learning. But the boundary blur of the depth map is still a serious problem. Research finds the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Xin Yang , Qingling Chang , Xinlin Liu , Yan Cui

Perceiving 3D objects from monocular inputs is crucial for robotic systems, given its economy compared to multi-sensor settings. It is notably difficult as a single image can not provide any clues for predicting absolute depth values.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Tai Wang , Jiangmiao Pang , Dahua Lin

UAVs have become an essential photogrammetric measurement as they are affordable, easily accessible and versatile. Aerial images captured from UAVs have applications in small and large scale texture mapping, 3D modelling, object detection…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Logambal Madhuanand , Francesco Nex , Michael Ying Yang

3D object detection is an important capability needed in various practical applications such as driver assistance systems. Monocular 3D detection, as a representative general setting among image-based approaches, provides a more economical…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Tai Wang , Xinge Zhu , Jiangmiao Pang , Dahua Lin

Spherical cameras capture scenes in a holistic manner and have been used for room layout estimation. Recently, with the availability of appropriate datasets, there has also been progress in depth estimation from a single omnidirectional…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Nikolaos Zioulis , Federico Alvarez , Dimitrios Zarpalas , Petros Daras

Traditional monocular depth estimation suffers from inherent ambiguity and visual nuisances. We demonstrate that language can enhance monocular depth estimation by providing an additional condition (rather than images alone) aligned with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Ziyao Zeng , Jingcheng Ni , Daniel Wang , Patrick Rim , Younjoon Chung , Fengyu Yang , Byung-Woo Hong , Alex Wong

Single-view depth prediction is a fundamental problem in computer vision. Recently, deep learning methods have led to significant progress, but such methods are limited by the available training data. Current datasets based on 3D sensors…

Computer Vision and Pattern Recognition · Computer Science 2018-11-29 Zhengqi Li , Noah Snavely

This paper tackles the challenges of self-supervised monocular depth estimation in indoor scenes caused by large rotation between frames and low texture. We ease the learning process by obtaining coarse camera poses from monocular sequences…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Chaoqiang Zhao , Matteo Poggi , Fabio Tosi , Lei Zhou , Qiyu Sun , Yang Tang , Stefano Mattoccia

Monocular 3D object detection has attracted widespread attention due to its potential to accurately obtain object 3D localization from a single image at a low cost. Depth estimation is an essential but challenging subtask of monocular 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Longfei Yan , Pei Yan , Shengzhou Xiong , Xuanyu Xiang , Yihua Tan

For the task of simultaneous monocular depth and visual odometry estimation, we propose learning self-supervised transformer-based models in two steps. Our first step consists in a generic pretraining to learn 3D geometry, using cross-view…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Boris Chidlovskii , Leonid Antsfeld

In monocular depth estimation, disturbances in the image context, like moving objects or reflecting materials, can easily lead to erroneous predictions. For that reason, uncertainty estimates for each pixel are necessary, in particular for…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Julia Hornauer , Vasileios Belagiannis

Monocular depth priors have been widely adopted by neural rendering in multi-view based tasks such as 3D reconstruction and novel view synthesis. However, due to the inconsistent prediction on each view, how to more effectively leverage…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Wenyuan Zhang , Yixiao Yang , Han Huang , Liang Han , Kanle Shi , Yu-Shen Liu , Zhizhong Han

Depth estimation is a challenging task of 3D reconstruction to enhance the accuracy sensing of environment awareness. This work brings a new solution with a set of improvements, which increase the quantitative and qualitative understanding…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Armin Masoumian , Hatem A. Rashwan , Saddam Abdulwahab , Julian Cristiano , Domenec Puig

Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained visual perception and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Hanxun Yu , Xuan Qu , Yuxin Wang , Jianke Zhu , Lei Ke

Monocular depth estimation in the wild inherently predicts depth up to an unknown scale. To resolve scale ambiguity issue, we present a learning algorithm that leverages monocular simultaneous localization and mapping (SLAM) with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Jaehoon Choi , Dongki Jung , Yonghan Lee , Deokhwa Kim , Dinesh Manocha , Donghwan Lee

Self-supervised learning of depth map prediction and motion estimation from monocular video sequences is of vital importance -- since it realizes a broad range of tasks in robotics and autonomous vehicles. A large number of research efforts…

Computer Vision and Pattern Recognition · Computer Science 2021-03-24 Ue-Hwan Kim , Jong-Hwan Kim

Supervised learning based methods for monocular depth estimation usually require large amounts of extensively annotated training data. In the case of aerial imagery, this ground truth is particularly difficult to acquire. Therefore, in this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Max Hermann , Boitumelo Ruf , Martin Weinmann , Stefan Hinz