English
Related papers

Related papers: Global Structure-from-Motion Meets Feedforward Rec…

200 papers

Structure from motion is the process of recovering information about cameras and 3D scene from a set of images. Generally, in a noise-free setting, all information can be uniquely recovered if enough images and image points are provided.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Martin Bråtelund

This position paper argues for the use of \emph{structured generative models} (SGMs) for the understanding of static scenes. This requires the reconstruction of a 3D scene from an input image (or a set of multi-view images), whereby the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Christopher K. I. Williams

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Congrong Xu , Huachen Gao , Xingyu Chen , Yuliang Xiu , Jun Gao , Anpei Chen

World-wide detailed 2D maps require enormous collective efforts. OpenStreetMap is the result of 11 million registered users manually annotating the GPS location of over 1.75 billion entries, including distinctive landmarks and common urban…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Matteo Toso , Stefano Fiorini , Stuart James , Alessio Del Bue

This work addresses the task of dense 3D reconstruction of a complex dynamic scene from images. The prevailing idea to solve this task is composed of a sequence of steps and is dependent on the success of several pipelines in its execution.…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Suryansh Kumar , Yuchao Dai , Hongdong Li

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Yu Sheng , Jiajun Deng , Xinran Zhang , Yu Zhang , Bei Hua , Yanyong Zhang , Jianmin Ji

We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing multi-agent reconstruction frameworks are often hindered by rigid prerequisites,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yihang Tao , Yu Guo , Zhengru Fang , Haonan An , Yuguang Fang

Structure from Motion (SfM) often fails to estimate accurate poses in environments that lack suitable visual features. In such cases, the quality of the final 3D mesh, which is contingent on the accuracy of those estimates, is reduced. One…

Computer Vision and Pattern Recognition · Computer Science 2021-04-08 Victor Amblard , Timothy P. Osedach , Arnaud Croux , Andrew Speck , John J. Leonard

Existing text-to-3D and image-to-3D models often struggle with complex scenes involving multiple objects and intricate interactions. Although some recent attempts have explored such compositional scenarios, they still require an extensive…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Yujia Hu , Songhua Liu , Xingyi Yang , Xinchao Wang

Accurate 3D reconstruction in visually-degraded underwater environments remains a formidable challenge. Single-modality approaches are insufficient: vision-based methods fail due to poor visibility and geometric constraints, while sonar is…

Robotics · Computer Science 2026-05-19 Lingpeng Chen , Jiakun Tang , Apple Pui-Yi Chui , Ziyang Hong , Junfeng Wu

Video frame prediction remains a fundamental challenge in computer vision with direct implications for autonomous systems, video compression, and media synthesis. We present FG-DFPN, a novel architecture that harnesses the synergy between…

Image and Video Processing · Electrical Eng. & Systems 2025-03-17 M. Akın Yılmaz , Ahmet Bilican , A. Murat Tekalp

We present a novel multi-altitude camera pose estimation system, addressing the challenges of robust and accurate localization across varied altitudes when only considering sparse image input. The system effectively handles diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Yaxuan Li , Yewei Huang , Bijay Gaudel , Hamidreza Jafarnejadsani , Brendan Englot

This paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel over-segmentation to the image, we model a generically dynamic (hence non-rigid) scene…

Computer Vision and Pattern Recognition · Computer Science 2017-12-21 Suryansh Kumar , Yuchao Dai , Hongdong Li

We present WorldMirror, an all-in-one, feed-forward model for versatile 3D geometric prediction tasks. Unlike existing methods constrained to image-only inputs or customized for a specific task, our framework flexibly integrates diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yifan Liu , Zhiyuan Min , Zhenwei Wang , Junta Wu , Tengfei Wang , Yixuan Yuan , Yawei Luo , Chunchao Guo

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models…

Robotics · Computer Science 2026-05-05 Sizhe Yang , Linning Xu , Hao Li , Juncheng Mu , Jia Zeng , Dahua Lin , Jiangmiao Pang

Incremental Structure from Motion (ISfM) has been widely used for UAV image orientation. Its efficiency, however, decreases dramatically due to the sequential constraint. Although the divide-and-conquer strategy has been utilized for…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 San Jiang , Qingquan Li , Wanshou Jiang , Wu Chen

Motion segmentation in dynamic scenes is highly challenging, as conventional methods heavily rely on estimating camera poses and point correspondences from inherently noisy motion cues. Existing statistical inference or iterative…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Xiankang He , Peile Lin , Ying Cui , Dongyan Guo , Chunhua Shen , Xiaoqin Zhang

We recover the underlying 3D structure from images of cartoons and anime depicting the same scene. This is an interesting problem domain because images in creative media are often depicted without explicit geometric consistency for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Ethan Weber , Riley Peterlinz , Rohan Mathur , Frederik Warburg , Alexei A. Efros , Angjoo Kanazawa

Two-view structure from motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM (vSLAM). Many existing end-to-end learning-based methods usually formulate it as a brute regression problem. However, the inadequate utilization of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Yuxi Xiao , Li Li , Xiaodi Li , Jian Yao

Our world is full of identical objects (\emphe.g., cans of coke, cars of same model). These duplicates, when seen together, provide additional and strong cues for us to effectively reason about 3D. Inspired by this observation, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Tianhang Cheng , Wei-Chiu Ma , Kaiyu Guan , Antonio Torralba , Shenlong Wang