English
Related papers

Related papers: Depth Anything 3: Recovering the Visual Space from…

200 papers

Perceiving 3D information is of paramount importance in many applications of computer vision. Recent advances in monocular depth estimation have shown that gaining such knowledge from a single camera input is possible by training deep…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Sai Shyam Chanduri , Zeeshan Khan Suri , Igor Vozniak , Christian Müller

We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual…

Visual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Siyan Dong , Shuzhe Wang , Shaohui Liu , Lulu Cai , Qingnan Fan , Juho Kannala , Yanchao Yang

Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu

Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yuecheng Liu , Junda Cheng , Longliang Liu , Wenjing Liao , Hanrui Cheng , Yuzhou Wang , Xin Yang

Recovering the scene depth from a single image is an ill-posed problem that requires additional priors, often referred to as monocular depth cues, to disambiguate different 3D interpretations. In recent works, those priors have been learned…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Lam Huynh , Phong Nguyen-Ha , Jiri Matas , Esa Rahtu , Janne Heikkila

3D reconstruction from 2D inputs, especially for non-rigid objects like humans, presents unique challenges due to the significant range of possible deformations. Traditional methods often struggle with non-rigid shapes, which require…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Fahd Alhamazani , Yu-Kun Lai , Paul L. Rosin

We present an algorithm for reconstructing dense, geometrically consistent depth for all pixels in a monocular video. We leverage a conventional structure-from-motion reconstruction to establish geometric constraints on pixels in the video.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-28 Xuan Luo , Jia-Bin Huang , Richard Szeliski , Kevin Matzen , Johannes Kopf

This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Zehan Wang , Siyu Chen , Lihe Yang , Jialei Wang , Ziang Zhang , Hengshuang Zhao , Zhou Zhao

3D object detection is an important capability needed in various practical applications such as driver assistance systems. Monocular 3D detection, as a representative general setting among image-based approaches, provides a more economical…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Tai Wang , Xinge Zhu , Jiangmiao Pang , Dahua Lin

One practical approach to infer 3D scene structure from a single image is to retrieve a closely matching 3D model from a database and align it with the object in the image. Existing methods rely on supervised training with images and pose…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Pattaramanee Arsomngern , Sasikarn Khwanmuang , Matthias Nießner , Supasorn Suwajanakorn

We introduce MapAnything, a unified transformer-based feed-forward model that ingests one or more images along with optional geometric inputs such as camera intrinsics, poses, depth, or partial reconstructions, and then directly regresses…

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Ruicheng Wang , Sicheng Xu , Cassie Dai , Jianfeng Xiang , Yu Deng , Xin Tong , Jiaolong Yang

Accurately estimating depth in 360-degree imagery is crucial for virtual reality, autonomous navigation, and immersive media applications. Existing depth estimation methods designed for perspective-view imagery fail when applied to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Ning-Hsu Wang , Yu-Lun Liu

We present Pow3r, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3r…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Wonbong Jang , Philippe Weinzaepfel , Vincent Leroy , Lourdes Agapito , Jerome Revaud

Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3 -- a 3D…

Robotics · Computer Science 2025-09-22 Quanhao Qian , Guoyang Zhao , Gongjie Zhang , Jiuniu Wang , Ran Xu , Junlong Gao , Deli Zhao

Monocular depth estimation has drawn widespread attention from the vision community due to its broad applications. In this paper, we propose a novel physics (geometry)-driven deep learning framework for monocular depth estimation by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Shuwei Shao , Zhongcai Pei , Weihai Chen , Xingming Wu , Zhengguo Li

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiahao Lu , Tianyu Huang , Peng Li , Zhiyang Dou , Cheng Lin , Zhiming Cui , Zhen Dong , Sai-Kit Yeung , Wenping Wang , Yuan Liu

We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences featuring substantial object deformation, large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ting-Hsuan Liao , Haowen Liu , Yiran Xu , Songwei Ge , Gengshan Yang , Jia-Bin Huang