English
Related papers

Related papers: SAB3R: Semantic-Augmented Backbone in 3D Reconstru…

200 papers

Creating 3D semantic reconstructions of environments is fundamental to many applications, especially when related to autonomous agent operation (e.g., goal-oriented navigation or object interaction and manipulation). Commonly, 3D semantic…

Robotics · Computer Science 2024-06-11 Jianhao Zheng , Daniel Barath , Marc Pollefeys , Iro Armeni

Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely on multi-view fusion of semantic embeddings into a 3D map,…

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Yushi Lan , Yihang Luo , Fangzhou Hong , Shangchen Zhou , Honghua Chen , Zhaoyang Lyu , Shuai Yang , Bo Dai , Chen Change Loy , Xingang Pan

The recent development in multimodal learning has greatly advanced the research in 3D scene understanding in various real-world tasks such as embodied AI. However, most existing studies are facing two common challenges: 1) they are short of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xueying Jiang , Lewei Lu , Ling Shao , Shijian Lu

Image matching is a key component of modern 3D vision algorithms, essential for accurate scene reconstruction and localization. MASt3R redefines image matching as a 3D task by leveraging DUSt3R and introducing a fast reciprocal matching…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Jingxing Li , Yongjae Lee , Abhay Kumar Yadav , Cheng Peng , Rama Chellappa , Deliang Fan

Monocular Simultaneous Localization and Mapping (SLAM) aims to estimate a robot's pose while simultaneously reconstructing an unknown 3D scene using a single camera. While existing monocular SLAM systems generate detailed 3D geometry…

Robotics · Computer Science 2025-11-27 Yuchen Zhou , Haihang Wu

Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems. To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Qi Xu , Dongxu Wei , Lingzhe Zhao , Wenpu Li , Zhangchi Huang , Shunping Ji , Peidong Liu

Sparse and feature SLAM methods provide robust camera pose estimation. However, they often fail to capture the level of detail required for inspection and scene awareness tasks. Conversely, dense SLAM approaches generate richer scene…

Robotics · Computer Science 2025-05-16 Maaz Qureshi , Alexander Werner , Zhenan Liu , Amir Khajepour , George Shaker , William Melek

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Qianqian Wang , Yifei Zhang , Aleksander Holynski , Alexei A. Efros , Angjoo Kanazawa

3D fragment reassembly aims to recover the rigid poses of unordered fragment point clouds or meshes in a common object coordinate system to reconstruct the complete shape. The problem becomes particularly challenging as the number of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Hanze Jia , Chunshi Wang , Yuxiao Yang , Zhonghua Jiang , Yawei Luo , Shuainan Ye , Tan Tang

Robots often rely on RGB images for tasks like manipulation and navigation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on…

Robotics · Computer Science 2025-09-11 Davide Allegro , Matteo Terreran , Stefano Ghidoni

Large-scale semantic mapping is crucial for outdoor autonomous agents to fulfill high-level tasks such as planning and navigation. This paper proposes a novel method for large-scale 3D semantic reconstruction through implicit…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Jianyuan Zhang , Zhiliu Yang , Meng Zhang

Localization is an essential task for mobile autonomous robotic systems that want to use pre-existing maps or create new ones in the context of SLAM. Today, many robotic platforms are equipped with high-accuracy 3D LiDAR sensors, which…

In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection. Maps can provide robust structural priors of the static environment, helping resolve…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yang Fu , Yuliang Zou , Hao Xiang , Xin Huang , Yijing Bai , Chen Song , Weijing Shi , Govind Thattai , Dragomir Anguelov , Mingxing Tan , Yingwei Li

Robust 3D geometry estimation from videos is critical for applications such as autonomous navigation, SLAM, and 3D scene reconstruction. Recent methods like DUSt3R demonstrate that regressing dense pointmaps from image pairs enables…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Xiaoshan Wu , Yifei Yu , Xiaoyang Lyu , Yihua Huang , Bo Wang , Baoheng Zhang , Zhongrui Wang , Xiaojuan Qi

Research into dynamic 3D scene understanding has primarily focused on short-term change tracking from dense observations, while little attention has been paid to long-term changes with sparse observations. We address this gap with MoRE, a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Liyuan Zhu , Shengyu Huang , Konrad Schindler , Iro Armeni

Large Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Ruikai Cui , Xibin Song , Weixuan Sun , Senbo Wang , Weizhe Liu , Shenzhou Chen , Taizhang Shang , Yang Li , Nick Barnes , Hongdong Li , Pan Ji

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jinjie Mai , Wenxuan Zhu , Haozhe Liu , Bing Li , Cheng Zheng , Jürgen Schmidhuber , Bernard Ghanem

The goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which involves: 1)…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Yuheng Lu , Chenfeng Xu , Xiaobao Wei , Xiaodong Xie , Masayoshi Tomizuka , Kurt Keutzer , Shanghang Zhang

3D panoptic segmentation is a challenging perception task, especially in autonomous driving. It aims to predict both semantic and instance annotations for 3D points in a scene. Although prior 3D panoptic segmentation approaches have…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Zihao Xiao , Longlong Jing , Shangxuan Wu , Alex Zihao Zhu , Jingwei Ji , Chiyu Max Jiang , Wei-Chih Hung , Thomas Funkhouser , Weicheng Kuo , Anelia Angelova , Yin Zhou , Shiwei Sheng