English
Related papers

Related papers: GeoFusionLRM: Geometry-Aware Self-Correction for C…

200 papers

Monocular depth estimation (MDE) is a fundamental yet inherently ill-posed task. Recent vision foundation models (VFMs), particularly DINO-based transformers, have significantly improved accuracy and generalization for dense prediction.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Gongshu Wang , Zhirui Wang , Kan Yang

Building good 3D maps is a challenging and expensive task, which requires high-quality sensors and careful, time-consuming scanning. We seek to reduce the cost of building good reconstructions by correcting views of existing low-quality…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Ştefan Săftescu , Paul Newman

The precise reconstruction of 3D objects from a single RGB image in complex scenes presents a critical challenge in virtual reality, autonomous driving, and robotics. Existing neural implicit 3D representation methods face significant…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Luoxi Zhang , Pragyan Shrestha , Yu Zhou , Chun Xie , Itaru Kitahara

Neural implicit surface reconstruction using volume rendering techniques has recently achieved significant advancements in creating high-fidelity surfaces from multiple 2D images. However, current methods primarily target scenes with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Lintao Xiang , Hongpei Zheng , Bailin Deng , Hujun Yin

Embodied AI training and evaluation require object-centric digital twin environments with accurate metric geometry and semantic grounding. Recent transformer-based feedforward reconstruction methods can efficiently predict global point…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Quanyun Wu , Kyle Gao , Daniel Long , David A. Clausi , Jonathan Li , Yuhao Chen

We present HI-SLAM2, a geometry-aware Gaussian SLAM system that achieves fast and accurate monocular scene reconstruction using only RGB input. Existing Neural SLAM or 3DGS-based SLAM methods often trade off between rendering quality and…

Robotics · Computer Science 2026-02-03 Wei Zhang , Qing Cheng , David Skuddis , Niclas Zeller , Daniel Cremers , Norbert Haala

Recently, 3D Gaussian Splatting has emerged as a prominent research direction owing to its ultrarapid training speed and high-fidelity rendering capabilities. However, the unstructured and irregular nature of Gaussian point clouds poses…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Xiao Ren , Yu Liu , Ning An , Jian Cheng , Xin Qiao , He Kong

3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learning approaches have shown promising results, most still rely…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Adeela Islam , Stefano Fiorini , Manuel Lecha , Theodore Tsesmelis , Stuart James , Pietro Morerio , Alessio Del Bue

Recently, deep learning-based 3D face reconstruction methods have demonstrated promising advancements in terms of quality and efficiency. Nevertheless, these techniques face challenges in effectively handling occluded scenes and fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dapeng Zhao

We introduce TransformerFusion, a transformer-based 3D scene reconstruction approach. From an input monocular RGB video, the video frames are processed by a transformer network that fuses the observations into a volumetric feature grid…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Aljaž Božič , Pablo Palafox , Justus Thies , Angela Dai , Matthias Nießner

Reconstructing real-world objects from multi-view images is essential for applications in 3D editing, AR/VR, and digital content creation. Existing methods typically prioritize either geometric accuracy (Multi-View Stereo) or photorealistic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zhejia Cai , Puhua Jiang , Shiwei Mao , Hongkun Cao , Ruqi Huang

We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene modeling and multi-scale semantic reasoning to enable high-fidelity extreme zoom-in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jiale Shi , Jiarui Hu , Zesong Yang , Kaixuan Luan , Hujun Bao , Zhaopeng Cui

3D generative modeling is accelerating as the technology allowing the capture of geometric data is developing. However, the acquired data is often inconsistent, resulting in unregistered meshes or point clouds. Many generative learning…

Computer Vision and Pattern Recognition · Computer Science 2023-06-29 Thomas Besnier , Sylvain Arguillère , Emery Pierson , Mohamed Daoudi

Geometric estimation is required for scene understanding and analysis in panoramic 360{\deg} images. Current methods usually predict a single feature, such as depth or surface normal. These methods can lack robustness, especially when…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kun Huang , Fang-Lue Zhang , Fangfang Zhang , Yu-Kun Lai , Paul L. Rosin , Neil A. Dodgson

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving deep spatial…

Recent advances in dense 3D reconstruction have led to significant progress, yet achieving accurate unified geometric prediction remains a major challenge. Most existing methods are limited to predicting a single geometry quantity from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xianze Fang , Jingnan Gao , Zhe Wang , Zhuo Chen , Xingyu Ren , Jiangjing Lyu , Qiaomu Ren , Zhonglei Yang , Xiaokang Yang , Yichao Yan , Chengfei Lyu

Generalizable NeRF aims to synthesize novel views for unseen scenes. Common practices involve constructing variance-based cost volumes for geometry reconstruction and encoding 3D descriptors for decoding novel views. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Tianqi Liu , Xinyi Ye , Min Shi , Zihao Huang , Zhiyu Pan , Zhan Peng , Zhiguo Cao

High-quality geometric diagram generation presents both a challenge and an opportunity: it demands strict spatial accuracy while offering well-defined constraints to guide generation. Inspired by recent advances in geometry problem solving…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Xiaojing Wei , Ting Zhang , Wei He , Jingdong Wang , Hua Huang

Reconstructing high-quality point clouds from images remains challenging in computer vision. Existing generative-model-based approaches, particularly diffusion-model approaches that directly learn the posterior, may suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Seunghyeok Shin , Dabin Kim , Hongki Lim

Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Shihua Zhang , Qiuhong Shen , Shizun Wang , Tianbo Pan , Xinchao Wang