English
Related papers

Related papers: Level-S$^2$fM: Structure from Motion on Neural Lev…

200 papers

3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advances in Multi-modal Large Language Models (MLLMs) have…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Beining Xu , Siting Zhu , Zhao Jin , Junxian Li , Hesheng Wang

Both self-supervised depth estimation and Structure-from-Motion (SfM) recover scene depth from RGB videos. Despite sharing a similar objective, the two approaches are disconnected. Prior works of self-supervision backpropagate losses…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Shengjie Zhu , Xiaoming Liu

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Nanxi Li , Xiang Wang , Yuanjie Chen , Haode Zhang , Hong Li , Yong-Lu Li

We present a novel neural surface reconstruction method called NeuralRoom for reconstructing room-sized indoor scenes directly from a set of 2D images. Recently, implicit neural representations have become a promising way to reconstruct…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Yusen Wang , Zongcheng Li , Yu Jiang , Kaixuan Zhou , Tuo Cao , Yanping Fu , Chunxia Xiao

3D Gaussian Splatting (3DGS) has demonstrated remarkable real-time performance in novel view synthesis, yet its effectiveness relies heavily on dense multi-view inputs with precisely known camera poses, which are rarely available in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Zongqi He , Hanmin Li , Kin-Chung Chan , Yushen Zuo , Hao Xie , Zhe Xiao , Jun Xiao , Kin-Man Lam

Structure from Motion (SfM) estimates camera poses and reconstructs point clouds, forming a foundation for various tasks. However, applying SfM to driving scenes captured by multi-camera systems presents significant difficulties, including…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Lingfeng Xuan , Chang Nie , Yiqing Xu , Zhe Liu , Yanzi Miao , Hesheng Wang

Neural signed distance functions (SDFs) have shown powerful ability in fitting the shape geometry. However, inferring continuous signed distance fields from discrete unoriented point clouds still remains a challenge. The neural network…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Shengtao Li , Ge Gao , Yudong Liu , Ming Gu , Yu-Shen Liu

Typical Structure-from-Motion (SfM) pipelines rely on finding correspondences across images, recovering the projective structure of the observed scene and upgrading it to a metric frame using camera self-calibration constraints. Solving…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Rui Gong , Danda Pani Paudel , Ajad Chhatkuli , Luc Van Gool

Recovering the 3D structure of the surrounding environment is an essential task in any vision-controlled Structure-from-Motion (SfM) scheme. This paper focuses on the theoretical properties of the SfM, known as the incremental active depth…

Robotics · Computer Science 2020-03-17 Romulo T. Rodrigues , Pedro Miraldo , Dimos V. Dimarogonas , A. Pedro Aguiar

As commonly used implicit geometry representations, the signed distance function (SDF) is limited to modeling watertight shapes, while the unsigned distance function (UDF) is capable of representing various surfaces. However, its inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Chuanxiang Yang , Yuanfeng Zhou , Guangshun Wei , Long Ma , Junhui Hou , Yuan Liu , Wenping Wang

We propose a new structure-from-motion framework to recover accurate camera poses and point clouds from unordered images. Traditional SfM systems typically rely on the successful detection of repeatable keypoints across multiple views as…

Computer Vision and Pattern Recognition · Computer Science 2023-06-28 Xingyi He , Jiaming Sun , Yifan Wang , Sida Peng , Qixing Huang , Hujun Bao , Xiaowei Zhou

Structure from Motion (SfM) refers to the problem of recovering both structure (i.e., 3D coordinates of points in the scene) and motion (i.e., camera matrices) starting from point correspondences in multiple images. It has attracted…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Federica Arrigoni

State-of-the-art neural implicit surface representations have achieved impressive results in indoor scene reconstruction by incorporating monocular geometric priors as additional supervision. However, we have observed that multi-view…

Computer Vision and Pattern Recognition · Computer Science 2024-05-02 Ziyi Chen , Xiaolong Wu , Yu Zhang

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Qitao Zhao , Amy Lin , Jeff Tan , Jason Y. Zhang , Deva Ramanan , Shubham Tulsiani

We consider the problem of simultaneously estimating a dense depth map and camera pose for a large set of images of an indoor scene. While classical SfM pipelines rely on a two-step approach where cameras are first estimated using a bundle…

Computer Vision and Pattern Recognition · Computer Science 2020-11-23 Benjamin Graham , David Novotny

Structure-from-Motion (SfM) is the task of estimating 3D structure and camera poses from images. We define Collaborative SfM (ColabSfM) as sharing distributed SfM reconstructions. Sharing maps requires estimating a joint reference frame,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Johan Edstedt , André Mateus , Alberto Jaenal

Neural implicit surface reconstruction using volume rendering techniques has recently achieved significant advancements in creating high-fidelity surfaces from multiple 2D images. However, current methods primarily target scenes with…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Lintao Xiang , Hongpei Zheng , Bailin Deng , Hujun Yin

Recent advances have enabled a single neural network to serve as an implicit scene representation, establishing the mapping function between spatial coordinates and scene properties. In this paper, we make a further step towards continual…

Computer Vision and Pattern Recognition · Computer Science 2021-10-05 Zike Yan , Yuxin Tian , Xuesong Shi , Ping Guo , Peng Wang , Hongbin Zha

Structure-from-Motion is a technology used to obtain scene structure through image collection, which is a fundamental problem in computer vision. For unordered Internet images, SfM is very slow due to the lack of prior knowledge about image…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Zhichao Ye , Chong Bao , Xin Zhou , Haomin Liu , Hujun Bao , Guofeng Zhang

Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Wang Zhao , Shaohui Liu , Hengkai Guo , Wenping Wang , Yong-Jin Liu