English
Related papers

Related papers: MoGe-2: Accurate Monocular Geometry with Metric Sc…

200 papers

Most 3D reconstruction methods may only recover scene properties up to a global scale ambiguity. We present a novel approach to single view metrology that can recover the absolute scale of a scene represented by 3D heights of objects or…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Rui Zhu , Xingyi Yang , Yannick Hold-Geoffroy , Federico Perazzi , Jonathan Eisenmann , Kalyan Sunkavalli , Manmohan Chandraker

Joint camera pose and dense geometry estimation from a set of images or a monocular video remains a challenging problem due to its computational complexity and inherent visual ambiguities. Most dense incremental reconstruction systems…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Kirill Mazur , Gwangbin Bae , Andrew J. Davison

Convolutional Neural Networks have demonstrated superior performance on single image depth estimation in recent years. These works usually use stacked spatial pooling or strided convolution to get high-level information which are common…

Computer Vision and Pattern Recognition · Computer Science 2018-09-05 Zhixiang Hao , Yu Li , Shaodi You , Feng Lu

Depth estimation is one of the key technologies for realizing 3D perception in unmanned systems. Monocular depth estimation has been widely researched because of its low-cost advantage, but the existing methods face the challenges of poor…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Chang Liu , Juan Li , Sheng Zhang , Chang Liu , Jie Li , Xu Zhang

How can we effectively utilise the 2D monocular image information for recovering the 6D pose (6-DoF) of the visual objects? Deep learning has shown to be effective for robust and real-time monocular pose estimation. Oftentimes, the network…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Di Wu , Yihao Chen , Xianbiao Qi , Yongjian Yu , Weixuan Chen , Rong Xiao

Estimating accurate, view-consistent geometry and camera poses from uncalibrated multi-view/video inputs remains challenging - especially at high spatial resolutions and over long sequences. We present DAGE, a dual-stream transformer whose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Tuan Duc Ngo , Jiahui Huang , Seoung Wug Oh , Kevin Blackburn-Matzen , Evangelos Kalogerakis , Chuang Gan , Joon-Young Lee

Depth estimation plays a pivotal role in advancing human-robot interactions, especially in indoor environments where accurate 3D scene reconstruction is essential for tasks like navigation and object handling. Monocular depth estimation,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Siddiqui Muhammad Yasir , Hyunsik Ahn

One of the key elements of reconstructing a 3D mesh from a monocular video is generating every frame's depth map. However, in the application of colonoscopy video reconstruction, producing good-quality depth estimation is challenging.…

Image and Video Processing · Electrical Eng. & Systems 2021-11-23 Yubo Zhang , Jan-Michael Frahm , Samuel Ehrenstein , Sarah K. McGill , Julian G. Rosenman , Shuxian Wang , Stephen M. Pizer

Accurately perceiving location and scene is crucial for autonomous driving and mobile robots. Recent advances in deep learning have made it possible to learn egomotion and depth from monocular images in a self-supervised manner, without…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Hao Qu , Lilian Zhang , Xiaoping Hu , Xiaofeng He , Xianfei Pan , Changhao Chen

Reconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA)…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Wei Yin , Chi Zhang , Hao Chen , Zhipeng Cai , Gang Yu , Kaixuan Wang , Xiaozhi Chen , Chunhua Shen

Due to the inherent ill-posed nature of 2D-3D projection, monocular 3D object detection lacks accurate depth recovery ability. Although the deep neural network (DNN) enables monocular depth-sensing from high-level learned features, the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Qing Lian , Peiliang Li , Xiaozhi Chen

Recovering 3D full-body human pose is a challenging problem with many applications. It has been successfully addressed by motion capture systems with body worn markers and multiple cameras. In this paper, we address the more challenging…

Computer Vision and Pattern Recognition · Computer Science 2018-03-12 Xiaowei Zhou , Menglong Zhu , Georgios Pavlakos , Spyridon Leonardos , Kostantinos G. Derpanis , Kostas Daniilidis

Single image depth estimation is a challenging problem. The current state-of-the-art method formulates the problem as that of ordinal regression. However, the formulation is not fully differentiable and depth maps are not generated in an…

Computer Vision and Pattern Recognition · Computer Science 2020-06-17 Kunal Swami , Prasanna Vishnu Bondada , Pankaj Kumar Bajpai

Estimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth…

Robotics · Computer Science 2022-11-29 Ruofeng Wei , Bin Li , Hangjie Mo , Fangxun Zhong , Yonghao Long , Qi Dou , Yun-Hui Liu , Dong Sun

Monocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as…

Computer Vision and Pattern Recognition · Computer Science 2022-06-30 Xuepeng Shi , Qi Ye , Xiaozhi Chen , Chuangrong Chen , Zhixiang Chen , Tae-Kyun Kim

Recent advances in discriminative and generative pretraining have yielded geometry estimation models with strong generalization capabilities. While discriminative monocular geometry estimation methods rely on large-scale fine-tuning data to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Yongtao Ge , Guangkai Xu , Zhiyue Zhao , Libo Sun , Zheng Huang , Yanlong Sun , Hao Chen , Chunhua Shen

Monocular depth estimation is scale-ambiguous, and thus requires scale supervision to produce metric predictions. Even so, the resulting models will be geometry-specific, with learned scales that cannot be directly transferred across…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Vitor Guizilini , Igor Vasiljevic , Dian Chen , Rares Ambrus , Adrien Gaidon

Accurate reconstruction of arbitrary-shaped long slender continuum bodies, such as guidewires, catheters and other soft continuum manipulators, is essential for accurate mechanical simulation. However, existing image-based reconstruction…

Robotics · Computer Science 2026-03-25 Yaozhi Zhang , Shun Yu , Yugang Zhang , Yang Liu

3D object detection is an important capability needed in various practical applications such as driver assistance systems. Monocular 3D detection, as a representative general setting among image-based approaches, provides a more economical…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Tai Wang , Xinge Zhu , Jiangmiao Pang , Dahua Lin

Monocular depth estimation is a highly challenging problem that is often addressed with deep neural networks. While these are able to use recognition of image features to predict reasonably looking depth maps the result often has low metric…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Patrik Persson , Linn Öström , Carl Olsson