English
Related papers

Related papers: LightSplat: Fast and Memory-Efficient Open-Vocabul…

200 papers

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Yu Sheng , Jiajun Deng , Xinran Zhang , Yu Zhang , Bei Hua , Yanyong Zhang , Jianmin Ji

We present SceneVGGT, a spatio-temporal 3D scene understanding framework that combines SLAM with semantic mapping for autonomous and assistive navigation. Built on VGGT, our method scales to long video streams via a sliding-window pipeline.…

We study open-world 3D scene understanding, a family of tasks that require agents to reason about their 3D environment with an open-set vocabulary and out-of-domain visual inputs - a critical skill for robots to operate in the unstructured…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Huy Ha , Shuran Song

Photorealistic 3D reconstruction of unstructured real-world scenes remains challenging due to complex illumination variations and transient occlusions. Existing methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS)…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Yuzhou Tang , Dejun Xu , Yongjie Hou , Zhenzhong Wang , Min Jiang

A main bottleneck of learning-based robotic scene understanding methods is the heavy reliance on extensive annotated training data, which often limits their generalization ability. In LiDAR panoptic segmentation, this challenge becomes even…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Ahmet Selim Çanakçı , Niclas Vödisch , Kürsat Petek , Wolfram Burgard , Abhinav Valada

In recent years, there has been a surge of interest in open-vocabulary 3D scene reconstruction facilitated by visual language models (VLMs), which showcase remarkable capabilities in open-set retrieval. However, existing methods face some…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Yinan Deng , Jiahui Wang , Jingyu Zhao , Jianyu Dou , Yi Yang , Yufeng Yue

Open-Vocabulary Semantic Segmentation (OVSS) has advanced with recent vision-language models (VLMs), enabling segmentation beyond predefined categories through various learning schemes. Notably, training-free methods offer scalable, easily…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Chanyoung Kim , Dayun Ju , Woojung Han , Ming-Hsuan Yang , Seong Jae Hwang

In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in the 3D data. However, existing methods lack explicit mechanisms to incorporate geometric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Diogo Lavado , Alessandra Micheletti , Clàudia Soares

Building models that can understand and reason about 3D scenes is difficult owing to the lack of data sources for 3D supervised training and large-scale training regimes. In this work we ask - How can the knowledge in a pre-trained language…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Shivam Chandhok

In orthodontic treatment, particularly within telemedicine contexts, observing patients' dental occlusion from multiple viewpoints facilitates timely clinical decision-making. Recent advances in 3D Gaussian Splatting (3DGS) have shown…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Yiyi Miao , Taoyu Wu , Tong Chen , Sihao Li , Ji Jiang , Youpeng Yang , Angelos Stefanidis , Limin Yu , Jionglong Su

Recently, high-fidelity scene reconstruction with an optimized 3D Gaussian splat representation has been introduced for novel view synthesis from sparse image sets. Making such representations suitable for applications like network…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Simon Niedermayr , Josef Stumpfegger , Rüdiger Westermann

This paper proposes Neural-MMGS, a novel neural 3DGS framework for multimodal large-scale scene reconstruction that fuses multiple sensing modalities in a per-gaussian compact, learnable embedding. While recent works focusing on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Sitian Shen , Georgi Pramatarov , Yifu Tao , Daniele De Martini

Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jiazhong Cen , Xudong Zhou , Jiemin Fang , Changsong Wen , Lingxi Xie , Xiaopeng Zhang , Wei Shen , Qi Tian

The creation of detailed 3D models is relevant for a wide range of applications such as navigation in three-dimensional space, construction planning or disaster assessment. However, the complex processing and long execution time for…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Max Hermann , Thomas Pollok , Daniel Brommer , Dominic Zahn

Recent advancements in 3D editing have highlighted the potential of text-driven methods in real-time, user-friendly AR/VR applications. However, current methods rely on 2D diffusion models without adequately considering multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Dong In Lee , Hyeongcheol Park , Jiyoung Seo , Eunbyung Park , Hyunje Park , Ha Dam Baek , Sangheon Shin , Sangmin Kim , Sangpil Kim

Reconstructing intricate, ever-changing environments remains a central ambition in computer vision, yet existing solutions often crumble before the complexity of real-world dynamics. We present DynaSplat, an approach that extends Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Junli Deng , Ping Shi , Qipei Li , Jinyang Guo

The task of 3D semantic scene completion using monocular cameras is gaining significant attention in the field of autonomous driving. This task aims to predict the occupancy status and semantic labels of each voxel in a 3D scene from…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Jiawei Yao , Jusheng Zhang , Xiaochao Pan , Tong Wu , Canran Xiao

Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Chen Shi , Shaoshuai Shi , Xiaoyang Lyu , Chunyang Liu , Kehua Sheng , Bo Zhang , Li Jiang

3D affordance reasoning, the task of associating human instructions with the functional regions of 3D objects, is a critical capability for embodied agents. Current methods based on 3D Gaussian Splatting (3DGS) are fundamentally limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Yuhui Zheng , Mingtao Feng , Guangming Shi

3D scene segmentation based on neural implicit representation has emerged recently with the advantage of training only on 2D supervision. However, existing approaches still requires expensive per-scene optimization that prohibits…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Hanlin Chen , Chen Li , Mengqi Guo , Zhiwen Yan , Gim Hee Lee