English
Related papers

Related papers: Metric3D: Towards Zero-shot Metric 3D Prediction f…

200 papers

Understanding articulated objects from monocular video is a crucial yet challenging task in robotics and digital twin creation. Existing methods often rely on complex multi-view setups, high-fidelity object scans, or fragile long-term point…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Arslan Artykov , Tom Ravaud , Corentin Sautier , Vincent Lepetit

3D object detection from monocular image(s) is a challenging and long-standing problem of computer vision. To combine information from different perspectives without troublesome 2D instance tracking, recent methods tend to aggregate…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Jianlin Liu , Zhuofei Huang , Dihe Huang , Shang Xu , Ying Chen , Yong Liu

Reconstructing 3D scenes from monocular surgical videos can enhance surgeon's perception and therefore plays a vital role in various computer-assisted surgery tasks. However, achieving scale-consistent reconstruction remains an open…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jiaxin Guo , Wenzhen Dong , Tianyu Huang , Hao Ding , Ziyi Wang , Haomin Kuang , Qi Dou , Yun-Hui Liu

3D human pose estimation is frequently seen as the task of estimating 3D poses relative to the root body joint. Alternatively, we propose a 3D human pose estimation method in camera coordinates, which allows effective combination of 2D…

Computer Vision and Pattern Recognition · Computer Science 2021-08-23 Diogo C Luvizon , Hedi Tabia , David Picard

Monocular depth estimation (MDE) has widely applicable but remains highly challenging due to the inherently ill-posed nature of reconstructing 3D scenes from single 2D images. Modern Vision Foundation Models (VFMs), pre-trained on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Gongshu Wang , Zhirui Wang , Kan Yang

Self-supervised monocular depth estimation has shown impressive results in static scenes. It relies on the multi-view consistency assumption for training networks, however, that is violated in dynamic object regions and occlusions.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Libo Sun , Jia-Wang Bian , Huangying Zhan , Wei Yin , Ian Reid , Chunhua Shen

In this paper, we present a system for incrementally reconstructing a dense 3D model of the geometry of an outdoor environment using a single monocular camera attached to a moving vehicle. Dense models provide a rich representation of the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Louis Gallagher , Varun Ravi Kumar , Senthil Yogamani , John B. McDonald

Regardless of the tremendous progress, a truly general purpose pipeline for Simultaneous Localization and Mapping (SLAM) remains a challenge. We investigate the reported failure of state of the art (SOTA) SLAM techniques on egocentric…

Computer Vision and Pattern Recognition · Computer Science 2018-11-20 Suvam Patra , Kartikeya Gupta , Faran Ahmad , Chetan Arora , Subhashis Banerjee

Current motion-based multiple object tracking (MOT) approaches rely heavily on Intersection-over-Union (IoU) for object association. Without using 3D features, they are ineffective in scenarios with occlusions or visually similar objects.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Milad Khanchi , Maria Amer , Charalambos Poullis

The detection of 3D objects through a single perspective camera is a challenging issue. The anchor-free and keypoint-based models receive increasing attention recently due to their effectiveness and simplicity. However, most of these…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Wei Chen , Jie Zhao , Wan-Lei Zhao , Song-Yuan Wu

Scene depth estimation from stereo and monocular imagery is critical for extracting 3D information for downstream tasks such as scene understanding. Recently, learning-based methods for depth estimation have received much attention due to…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Zhaoshuo Li , Nathan Drenkow , Hao Ding , Andy S. Ding , Alexander Lu , Francis X. Creighton , Russell H. Taylor , Mathias Unberath

Pre-trained Vision-Language Models (VLMs), such as CLIP, have shown enhanced performance across a range of tasks that involve the integration of visual and linguistic modalities. When CLIP is used for depth estimation tasks, the patches,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Xueting Hu , Ce Zhang , Yi Zhang , Bowen Hai , Ke Yu , Zhihai He

Unsupervised methods have showed promising results on monocular depth estimation. However, the training data must be captured in scenes without moving objects. To push the envelope of accuracy, recent methods tend to increase their model…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Tak-Wai Hui

Single-view depth prediction is a fundamental problem in computer vision. Recently, deep learning methods have led to significant progress, but such methods are limited by the available training data. Current datasets based on 3D sensors…

Computer Vision and Pattern Recognition · Computer Science 2018-11-29 Zhengqi Li , Noah Snavely

Multi-Object Tracking (MOT) remains a vital component of intelligent video analysis, which aims to locate targets and maintain a consistent identity for each target throughout a video sequence. Existing works usually learn a discriminative…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Yizhe Li , Sanping Zhou , Zheng Qin , Le Wang , Jinjun Wang , Nanning Zheng

Solving the challenging problem of 3D object reconstruction from a single image appropriately gives existing technologies the ability to perform with a single monocular camera rather than requiring depth sensors. In recent years, thanks to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Guiju Ping , Mahdi Abolfazli Esfahani , Han Wang

Deep metric learning aims to learn features relying on the consistency or divergence of class labels. However, in monocular depth estimation, the absence of a natural definition of class poses challenges in the leveraging of deep metric…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Chunpu Liu , Guanglei Yang , Wangmeng Zuo , Tianyi Zan

Existing methods for scale-invariant monocular depth estimation (SI MDE) often struggle due to the complexity of the task, and limited and non-diverse datasets, hindering generalizability in real-world scenarios. This is while…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 S. Mahdi H. Miangoleh , Mahesh Reddy , Yağız Aksoy

For the task of simultaneous monocular depth and visual odometry estimation, we propose learning self-supervised transformer-based models in two steps. Our first step consists in a generic pretraining to learn 3D geometry, using cross-view…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Boris Chidlovskii , Leonid Antsfeld

Monocular depth estimation (MDE), inferring pixel-level depths in single RGB images from a monocular camera, plays a crucial and pivotal role in a variety of AI applications demanding a three-dimensional (3D) topographical scene. In the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Mingyu Sung , Hyeonmin Choe , Il-Min Kim , Sangseok Yun , Jae Mo Kang