English
Related papers

Related papers: GaussianOcc3D: A Gaussian-Based Adaptive Multi-mod…

200 papers

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Xueyan Zou , Yuchen Song , Ri-Zhao Qiu , Xuanbin Peng , Jianglong Ye , Sifei Liu , Xiaolong Wang

Understanding and reconstructing the 3D world through omnidirectional perception is an inevitable trend in the development of autonomous agents and embodied intelligence. However, existing 3D occupancy prediction methods are constrained by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Mengfei Duan , Hao Shi , Fei Teng , Guoqiang Zhao , Yuheng Zhang , Zhiyong Li , Kailun Yang

Monocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image. In this paper, we strive to boost currently underperforming monocular…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Weijia Zhang , Dongnan Liu , Chao Ma , Weidong Cai

Open-vocabulary querying in 3D space is challenging but essential for scene understanding tasks such as object localization and segmentation. Language-embedded scene representations have made progress by incorporating language features into…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Jin-Chuan Shi , Miao Wang , Hao-Bin Duan , Shao-Hua Guan

In autonomous vehicles, understanding the surrounding 3D environment of the ego vehicle in real-time is essential. A compact way to represent scenes while encoding geometric distances and semantic object information is via 3D semantic…

Robotics · Computer Science 2024-05-21 Samuel Sze , Lars Kunze

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse conditions. Thus,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lianqing Zheng , Jianan Liu , Runwei Guan , Long Yang , Shouyi Lu , Yuanzhe Li , Xiaokai Bai , Jie Bai , Zhixiong Ma , Hui-Liang Shen , Xichan Zhu

Comprehensive modeling of the surrounding 3D world is key to the success of autonomous driving. However, existing perception tasks like object detection, road structure segmentation, depth & elevation estimation, and open-set object…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yuqi Wang , Yuntao Chen , Xingyu Liao , Lue Fan , Zhaoxiang Zhang

3D semantic occupancy prediction is crucial for autonomous driving, providing a dense, semantically rich environmental representation. However, existing methods focus on in-distribution scenes, making them susceptible to Out-of-Distribution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yuheng Zhang , Mengfei Duan , Kunyu Peng , Yuhang Wang , Ruiping Liu , Fei Teng , Kai Luo , Zhiyong Li , Kailun Yang

Understanding 3D scenes semantically and spatially is crucial for the safe navigation of robots and autonomous vehicles, aiding obstacle avoidance and accurate trajectory planning. Camera-based 3D semantic occupancy prediction, which infers…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Junsu Kim , Junhee Lee , Ukcheol Shin , Jean Oh , Kyungdon Joo

3D Gaussian splatting (3DGS) has recently emerged as an alternative representation that leverages a 3D Gaussian-based representation and introduces an approximated volumetric rendering, achieving very fast rendering speed and promising…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Joo Chan Lee , Daniel Rho , Xiangyu Sun , Jong Hwan Ko , Eunbyung Park

Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction from monocular or…

Robotics · Computer Science 2026-05-01 Zeyu Jiang , Changqing Zhou , Xingxing Zuo , Changhao Chen

The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras provide rich texture information and LiDAR offers precise…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Zitian Wang , Zehao Huang , Yulu Gao , Naiyan Wang , Si Liu

We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Dubing Chen , Huan Zheng , Jin Fang , Xingping Dong , Xianfei Li , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Pin Tang , Zhongdao Wang , Guoqing Wang , Jilai Zheng , Xiangxuan Ren , Bailan Feng , Chao Ma

Egocentric scenes exhibit frequent occlusions, varied viewpoints, and dynamic interactions compared to typical scene understanding tasks. Occlusions and varied viewpoints can lead to multi-view semantic inconsistencies, while dynamic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Di Li , Jie Feng , Jiahao Chen , Weisheng Dong , Guanbin Li , Guangming Shi , Licheng Jiao

Occupancy prediction plays a pivotal role in autonomous driving. Previous methods typically construct dense 3D volumes, neglecting the inherent sparsity of the scene and suffering from high computational costs. To bridge the gap, we…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Haisong Liu , Yang Chen , Haiguang Wang , Zetong Yang , Tianyu Li , Jia Zeng , Li Chen , Hongyang Li , Limin Wang

We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of free-form language queries. This is a challenging problem…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Antonin Vobecky , Oriane Siméoni , David Hurych , Spyros Gidaris , Andrei Bursuc , Patrick Pérez , Josef Sivic

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Yubo Cui , Zhiheng Li , Jiaqiang Wang , Zheng Fang

Image matching is a fundamental and critical task in various visual applications, such as Simultaneous Localization and Mapping (SLAM) and image retrieval, which require accurate pose estimation. However, most existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Miao Fan , Mingrui Chen , Chen Hu , Shuchang Zhou

In recent years, 3D Gaussian splatting (3D-GS) has emerged as a novel scene representation approach. However, existing vision-only 3D-GS methods often rely on hand-crafted heuristics for point-cloud densification and face challenges in…

Robotics · Computer Science 2025-01-16 Sheng Hong , Chunran Zheng , Yishu Shen , Changze Li , Fu Zhang , Tong Qin , Shaojie Shen