English
Related papers

Related papers: Fin3R: Fine-tuning Feed-forward 3D Reconstruction …

200 papers

Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed-forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zihao Zhu , Wenyuan Zhao , Nuo Chen , Chao Tian , Zhiwen Fan

Feed-forward 3D reconstruction models such as DUSt3R, VGGT, and Depth Anything 3 (DA3) are transformer-based foundation models that infer camera geometry and dense scene structure in a single forward pass. Trained at scale in a supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jelena Bratulić , Sudhanshu Mittal , Thomas Brox , Christian Rupprecht

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Congrong Xu , Huachen Gao , Xingyu Chen , Yuliang Xiu , Jun Gao , Anpei Chen

Recent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

We present Pow3r, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3r…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Wonbong Jang , Philippe Weinzaepfel , Vincent Leroy , Lourdes Agapito , Jerome Revaud

Multimodal 3D object detection has garnered considerable interest in autonomous driving. However, multimodal detectors suffer from dimension mismatches that derive from fusing 3D points with 2D pixels coarsely, which leads to sub-optimal…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Guoxin Zhang , Ziying Song , Lin Liu , Zhonghong Ou

Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video inputs with hundreds or thousands of frames remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zecheng Tang , Jiaye Fu , Qiankun Gao , Haijie Li , Yanmin Wu , Jiaqi Zhang , Siwei Ma , Jian Zhang

3D super-resolution (3DSR) aims to reconstruct high-resolution (HR) 3D scenes from low-resolution (LR) multi-view images. Existing methods rely on dense LR inputs and per-scene optimization, which restricts the high-frequency priors for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xiang Feng , Xiangbo Wang , Tieshi Zhong , Chengkai Wang , Yiting Zhao , Tianxiang Xu , Zhenzhong Kuang , Feiwei Qin , Xuefei Yin , Yanming Zhu

Recent advances in 2D-to-3D perception have enabled the recovery of 3D scene semantics from unposed images. However, prevailing methods often suffer from limited generalization, reliance on per-scene optimization, and semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jie Hu , Shizun Wang , Xinchao Wang

Online monocular 3D reconstruction enables dense scene recovery from streaming video but remains fundamentally limited by the stability-adaptation dilemma: the reconstruction model must rapidly incorporate novel viewpoints while preserving…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Lanbo Xu , Liang Guo , Caigui Jiang , Cheng Wang

We introduce $\pi^3$, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yifan Wang , Jianjun Zhou , Haoyi Zhu , Wenzheng Chang , Yang Zhou , Zizun Li , Junyi Chen , Jiangmiao Pang , Chunhua Shen , Tong He

Feed-forward foundation models for multi-view 3-dimensional (3D) reconstruction have been trained on large-scale datasets of perspective images; when tested on wide field-of-view images, e.g., from a fisheye camera, their performance…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruxiao Duan , Erin Hong , Dongxu Zhao , Eric Turner , Alex Wong , Yunwen Zhou

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiahao Lu , Tianyu Huang , Peng Li , Zhiyang Dou , Cheng Lin , Zhiming Cui , Zhen Dong , Sai-Kit Yeung , Wenping Wang , Yuan Liu

Image stitching aim to align two images taken from different viewpoints into one seamless, wider image. However, when the 3D scene contains depth variations and the camera baseline is significant, noticeable parallax occurs-meaning the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Muhua Zhu , Xinhao Jin , Chengbo Wang , Yongcong Zhang , Yifei Xue , Tie Ji , Yizhen Lao

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulation learns a global,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Weirong Chen , Chuanxia Zheng , Ganlin Zhang , Andrea Vedaldi , Daniel Cremers

We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT, and AnySplat typically predict per-pixel point maps or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Chensheng Peng , Quentin Herau , Jiezhi Yang , Yichen Xie , Yihan Hu , Wenzhao Zheng , Matthew Strong , Masayoshi Tomizuka , Wei Zhan

Despite recent advances in feed-forward 3D Gaussian Splatting, generalizable 3D reconstruction remains challenging, particularly in multi-view correspondence modeling. Existing approaches face a fundamental trade-off: explicit methods…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Heng Jia , Linchao Zhu , Na Zhao

We present MonoPSR, a monocular 3D object detection method that leverages proposals and shape reconstruction. First, using the fundamental relations of a pinhole camera model, detections from a mature 2D object detector are used to generate…

Computer Vision and Pattern Recognition · Computer Science 2019-04-04 Jason Ku , Alex D. Pon , Steven L. Waslander

Dietary assessment is a key contributor to monitoring health status. Existing self-report methods are tedious and time-consuming with substantial biases and errors. Image-based food portion estimation aims to estimate food energy values…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Zeman Shao , Gautham Vinod , Jiangpeng He , Fengqing Zhu

We propose a simple yet effective approach to enhance the performance of feed-forward 3D reconstruction models. Existing methods often struggle near depth discontinuities, where standard regression losses encourage spatial averaging and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Zichen Wang , Ang Cao , Liam J. Wang , Jeong Joon Park