English
Related papers

Related papers: Depth Anything 3: Recovering the Visual Space from…

200 papers

This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Lihe Yang , Bingyi Kang , Zilong Huang , Zhen Zhao , Xiaogang Xu , Jiashi Feng , Hengshuang Zhao

Panoramic depth estimation provides a comprehensive solution for capturing complete $360^\circ$ environmental structural information, offering significant benefits for robotics and AR/VR applications. However, while extensively studied in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Hualie Jiang , Ziyang Song , Zhiqiang Lou , Rui Xu , Minglang Tan

Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jelena Bratulić , Sudhanshu Mittal , Thomas Brox , Christian Rupprecht

Monocular depth estimation aims to recover the depth information of 3D scenes from 2D images. Recent work has made significant progress, but its reliance on large-scale datasets and complex decoders has limited its efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Zeyu Ren , Zeyu Zhang , Wukai Li , Qingxiang Liu , Hao Tang

We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While depth and normal are geometrically related and highly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Mu Hu , Wei Yin , Chi Zhang , Zhipeng Cai , Xiaoxiao Long , Kaixuan Wang , Hao Chen , Gang Yu , Chunhua Shen , Shaojie Shen

We introduce $\pi^3$, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yifan Wang , Jianjun Zhou , Haoyi Zhu , Wenzheng Chang , Yang Zhou , Zizun Li , Junyi Chen , Jiangmiao Pang , Chunhua Shen , Tong He

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA)…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Wei Yin , Chi Zhang , Hao Chen , Zhipeng Cai , Gang Yu , Kaixuan Wang , Xiaozhi Chen , Chunhua Shen

Solving depth estimation with monocular cameras enables the possibility of widespread use of cameras as low-cost depth estimation sensors in applications such as autonomous driving and robotics. However, learning such a scalable depth…

Computer Vision and Pattern Recognition · Computer Science 2020-07-30 Bin Cheng , Inderjot Singh Saggu , Raunak Shah , Gaurav Bansal , Dinesh Bharadia

Monocular depth estimation remains challenging, as foundation models such as Depth Anything V2 (DA-V2) struggle with real-world images that are far from the training distribution. We introduce Re-Depth Anything, a test-time self-supervision…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ananta R. Bhattarai , Helge Rhodin

Although considerable advancements have been attained in self-supervised depth estimation from monocular videos, most existing methods often treat all objects in a video as static entities, which however violates the dynamic nature of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Xiuzhe Wu , Xiaoyang Lyu , Qihao Huang , Yong Liu , Yang Wu , Ying Shan , Xiaojuan Qi

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chi Zhang , Wei Yin , Gang Yu , Zhibin Wang , Tao Chen , Bin Fu , Joey Tianyi Zhou , Chunhua Shen

Recently, transformer-based methods have shown exceptional performance in monocular 3D object detection, which can predict 3D attributes from a single 2D image. These methods typically use visual and depth representations to generate query…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Xuan He , Jin Yuan , Kailun Yang , Zhenchao Zeng , Zhiyong Li

While recent depth foundation models exhibit strong zero-shot generalization, achieving accurate metric depth across diverse camera types-particularly those with large fields of view (FoV) such as fisheye and 360-degree cameras-remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yuliang Guo , Sparsh Garg , S. Mahdi H. Miangoleh , Xinyu Huang , Liu Ren

Recovering dense 3D geometry from unposed images remains a foundational challenge in computer vision. Current state-of-the-art models are predominantly trained on perspective datasets, which implicitly constrains them to a standard pinhole…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Namitha Guruprasad , Abhay Yadav , Cheng Peng , Rama Chellappa

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D vision. However, due…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Haodong Li , Wangguangdong Zheng , Jing He , Yuhao Liu , Xin Lin , Xin Yang , Ying-Cong Chen , Chunchao Guo

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

Multi-view stereo reconstruction (MVS) in the wild requires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangulate…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Shuzhe Wang , Vincent Leroy , Yohann Cabon , Boris Chidlovskii , Jerome Revaud

Single-view depth estimation (SVDE) plays a crucial role in scene understanding for AR applications, 3D modeling, and robotics, providing the geometry of a scene based on a single image. Recent works have shown that a successful solution…

Computer Vision and Pattern Recognition · Computer Science 2021-02-11 Mikhail Romanov , Nikolay Patatkin , Anna Vorontsova , Sergey Nikolenko , Anton Konushin , Dmitry Senyushkin

Monocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in terms of evaluation metrics such as the pixel-wise relative error, most methods neglect the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Wei Yin , Yifan Liu , Chunhua Shen

Despite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Long Mai , Simon Chen , Chunhua Shen
‹ Prev 1 2 3 10 Next ›