English
Related papers

Related papers: UniQueR: Unified Query-based Feedforward 3D Recons…

200 papers

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Congrong Xu , Huachen Gao , Xingyu Chen , Yuliang Xiu , Jun Gao , Anpei Chen

Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed-forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zihao Zhu , Wenyuan Zhao , Nuo Chen , Chao Tian , Zhiwen Fan

Representing 3D scenes from multiview images is a core challenge in computer vision and graphics, which requires both precise rendering and accurate reconstruction. Recently, 3D Gaussian Splatting (3DGS) has garnered significant attention…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 You Shen , Zhipeng Zhang , Xinyang Li , Yansong Qu , Yu Lin , Shengchuan Zhang , Liujuan Cao

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization using appearance embeddings or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Vinayak Gupta , Chih-Hao Lin , Shenlong Wang , Anand Bhattad , Jia-Bin Huang

We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap of all input images to a reference frame coordinate system,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Weining Ren , Hongjun Wang , Xiao Tan , Kai Han

Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jelena Bratulić , Sudhanshu Mittal , Thomas Brox , Christian Rupprecht

We propose a fast and accurate surface reconstruction algorithm for unorganized point clouds using an implicit representation. Recent learning methods are either single-object representations with small neural models that allow for high…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Siddhant Ranade , Gonçalo Dias Pais , Ross Tyler Whitaker , Jacinto C. Nascimento , Pedro Miraldo , Srikumar Ramalingam

Recently, 3D Gaussian Splatting (3DGS) has attracted widespread attention due to its high-quality rendering, and ultra-fast training and rendering speed. However, due to the unstructured and irregular nature of Gaussian point clouds, it is…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Danpeng Chen , Hai Li , Weicai Ye , Yifan Wang , Weijian Xie , Shangjin Zhai , Nan Wang , Haomin Liu , Hujun Bao , Guofeng Zhang

We present Any4D, a scalable multi-view transformer for metric-scale, dense feed-forward 4D reconstruction. Any4D directly generates per-pixel motion and geometry predictions for N frames, in contrast to prior work that typically focuses on…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Jay Karhade , Nikhil Keetha , Yuchen Zhang , Tanisha Gupta , Akash Sharma , Sebastian Scherer , Deva Ramanan

The scale diversity of point cloud data presents significant challenges in developing unified representation learning techniques for 3D vision. Currently, there are few unified 3D models, and no existing pre-training method is equally…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Ziyi Wang , Yanran Zhang , Jie Zhou , Jiwen Lu

3D scene reconstruction is essential for applications in virtual reality, robotics, and autonomous driving, enabling machines to understand and interact with complex environments. Traditional 3D Gaussian Splatting techniques rely on images…

Graphics · Computer Science 2025-03-04 Changlin Song , Jiaqi Wang , Liyun Zhu , He Weng

Feed-forward 3D Gaussian Splatting methods enable single-pass reconstruction and real-time rendering. However, they typically adopt rigid pixel-to-Gaussian or voxel-to-Gaussian pipelines that uniformly allocate Gaussians, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Injae Kim , Chaehyeon Kim , Minseong Bae , Minseok Joo , Hyunwoo J. Kim

Recent advances in dense 3D reconstruction have led to significant progress, yet achieving accurate unified geometric prediction remains a major challenge. Most existing methods are limited to predicting a single geometry quantity from…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Xianze Fang , Jingnan Gao , Zhe Wang , Zhuo Chen , Xingyu Ren , Jiangjing Lyu , Qiaomu Ren , Zhonglei Yang , Xiaokang Yang , Yichao Yan , Chengfei Lyu

In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no ground-truth annotations and no pretrained priors. During training, NAS3R reconstructs 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ranran Huang , Weixun Luo , Ye Mao , Krystian Mikolajczyk

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Weining Ren , Xiao Tan , Kai Han

We introduce GeoGS3D, a novel two-stage framework for reconstructing detailed 3D objects from single-view images. Inspired by the success of pre-trained 2D diffusion models, our method incorporates an orthogonal plane decomposition…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Qijun Feng , Zhen Xing , Zuxuan Wu , Yu-Gang Jiang

We present Implicit-Scale 3D Reconstruction from Monocular Multi-Food Images, a benchmark dataset designed to advance geometry-based food portion estimation in realistic dining scenarios. Existing dietary assessment methods largely rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuhao Chen , Gautham Vinod , Siddeshwar Raghavan , Talha Ibn Mahmud , Bruce Coburn , Jinge Ma , Fengqing Zhu , Jiangpeng He

Feed-forward multi-frame 3D reconstruction models often degrade on videos with object motion. Global-reference becomes ambiguous under multiple motions, while the local pointmap relies heavily on estimated relative poses and can drift,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Xingyu Miao , Weiguang Zhao , Tao Lu , Linning Xu , Mulin Yu , Yang Long , Jiangmiao Pang , Junting Dong

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulation learns a global,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Weirong Chen , Chuanxia Zheng , Ganlin Zhang , Andrea Vedaldi , Daniel Cremers

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction…