English
Related papers

Related papers: MapAnything: Universal Feed-Forward Metric 3D Reco…

200 papers

Depth maps are widely used in feed-forward 3D Gaussian Splatting (3DGS) pipelines by unprojecting them into 3D point clouds for novel view synthesis. This approach offers advantages such as efficient training, the use of known camera poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Duochao Shi , Weijie Wang , Donny Y. Chen , Zeyu Zhang , Jia-Wang Bian , Bohan Zhuang , Chunhua Shen

Most 3D reconstruction methods may only recover scene properties up to a global scale ambiguity. We present a novel approach to single view metrology that can recover the absolute scale of a scene represented by 3D heights of objects or…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Rui Zhu , Xingyi Yang , Yannick Hold-Geoffroy , Federico Perazzi , Jonathan Eisenmann , Kalyan Sunkavalli , Manmohan Chandraker

We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual…

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in spatially-aware tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Haiyi Qiu , Kaihang Pan , Jiacheng Li , Juncheng Li , Siliang Tang , Yueting Zhuang

We present a simple yet effective general-purpose framework for modeling 3D shapes by leveraging recent advances in 2D image generation using CNNs. Using just a single depth image of the object, we can output a dense multi-view depth map…

Computer Vision and Pattern Recognition · Computer Science 2020-09-08 Kamal Gupta , Susmija Jabbireddy , Ketul Shah , Abhinav Shrivastava , Matthias Zwicker

We propose a novel framework for comprehensive indoor 3D reconstruction using Gaussian representations, called OmniIndoor3D. This framework enables accurate appearance, geometry, and panoptic reconstruction of diverse indoor scenes captured…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Xiaobao Wei , Xiaoan Zhang , Hao Wang , Qingpo Wuwu , Ming Lu , Wenzhao Zheng , Shanghang Zhang

Despite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Long Mai , Simon Chen , Chunhua Shen

It has long been an ill-posed problem to predict absolute depth maps from single images in real (unseen) indoor scenes. We observe that it is essentially due to not only the scale-ambiguous problem but also the focal-ambiguous problem that…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Chengrui Wei , Meng Yang , Lei He , Nanning Zheng

We introduce AnySplat, a feed forward network for novel view synthesis from uncalibrated image collections. In contrast to traditional neural rendering pipelines that demand known camera poses and per scene optimization, or recent feed…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Lihan Jiang , Yucheng Mao , Linning Xu , Tao Lu , Kerui Ren , Yichen Jin , Xudong Xu , Mulin Yu , Jiangmiao Pang , Feng Zhao , Dahua Lin , Bo Dai

Monocular depth estimation is crucial for tracking and reconstruction algorithms, particularly in the context of surgical videos. However, the inherent challenges in directly obtaining ground truth depth maps during surgery render…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ange Lou , Yamin Li , Yike Zhang , Jack Noble

This paper presents a novel approach to reconstruct complete 3D deformable models over time by a single depth camera. These are the steps employed for deforming objects from single depth camera. The partial surfaces reconstructed from…

Computer Vision and Pattern Recognition · Computer Science 2017-08-31 Vamshhi Pavan Kumar Varma Vegeshna

3D reconstruction from a single-RGB image in unconstrained real-world scenarios presents numerous challenges due to the inherent diversity and complexity of objects and environments. In this paper, we introduce Anything-3D, a methodical…

Computer Vision and Pattern Recognition · Computer Science 2023-04-21 Qiuhong Shen , Xingyi Yang , Xinchao Wang

3D scene reconstruction is fundamental for spatial intelligence applications such as AR, robotics, and digital twins. Traditional multi-view stereo struggles with sparse viewpoints or low-texture regions, while neural rendering approaches,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Jiaqi Yao , Zhongmiao Yan , Jingyi Xu , Songpengcheng Xia , Yan Xiang , Ling Pei

Feed-forward 3D reconstruction methods aim to predict the 3D structure of a scene directly from input images, providing a faster alternative to per-scene optimization approaches. Significant progress has been made in single-view and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Sam Bahrami , Dylan Campbell

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, existing active…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Tianling Xu , Shengzhe Gan , Leslie Gu , Yuelei Li , Fangneng Zhan , Hanspeter Pfister

We introduce Intrinsic Image Fusion, a method that reconstructs high-quality physically based materials from multi-view images. Material reconstruction is highly underconstrained and typically relies on analysis-by-synthesis, which requires…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Peter Kocsis , Lukas Höllein , Matthias Nießner

Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Wonseok Roh , Hwanhee Jung , Jong Wook Kim , Seunggwan Lee , Innfarn Yoo , Andreas Lugmayr , Seunggeun Chi , Karthik Ramani , Sangpil Kim

Single-image 3D reconstruction is a research challenge focused on predicting 3D object shapes from single-view images. This task requires significant data acquisition to predict both visible and occluded portions of the shape. Furthermore,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Sanchar Palit , Sandika Biswas

Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, existing models have predominantly been developed and limited to be task-specific. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Sebastian Koch , Johanna Wald , Hidenobu Matsuki , Pedro Hermosilla , Timo Ropinski , Federico Tombari

We present FaceLift, a novel feed-forward approach for generalizable high-quality 360-degree 3D head reconstruction from a single image. Our pipeline first employs a multi-view latent diffusion model to generate consistent side and back…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Weijie Lyu , Yi Zhou , Ming-Hsuan Yang , Zhixin Shu