English
Related papers

Related papers: Lightweight Spatial Embedding for Vision-based 3D …

200 papers

We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Dubing Chen , Huan Zheng , Jin Fang , Xingping Dong , Xianfei Li , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen

The 3D occupancy prediction task has witnessed remarkable progress in recent years, playing a crucial role in vision-based autonomous driving systems. While traditional methods are limited to fixed semantic categories, recent approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Chi Yan , Dan Xu

Determining accurate bird's eye view (BEV) positions of objects and tracks in a scene is vital for various perception tasks including object interactions mapping, scenario extraction etc., however, the level of supervision required to…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 Paridhi Singh , Gaurav Singh , Arun Kumar

Understanding and modeling the 3D scene from a single image is a practical problem. A recent advance proposes a panoptic 3D scene reconstruction task that performs both 3D reconstruction and 3D panoptic segmentation from a single image.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Tao Chu , Pan Zhang , Qiong Liu , Jiaqi Wang

Accurate and reliable spatial and motion information plays a pivotal role in autonomous driving systems. However, object-level perception models struggle with handling open scenario categories and lack precise intrinsic geometry. On the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Kangan Qian , Jinyu Miao , Ziang Luo , Zheng Fu , and Jinchen Li , Yining Shi , Yunlong Wang , Kun Jiang , Mengmeng Yang , Diange Yang

Detecting diverse objects within complex indoor 3D point clouds presents significant challenges for robotic perception, particularly with varied object shapes, clutter, and the co-existence of static and dynamic elements where traditional…

Robotics · Computer Science 2025-07-24 Haichuan Li , Changda Tian , Panos Trahanias , Tomi Westerlund

3D semantic occupancy has rapidly become a research focus in the fields of robotics and autonomous driving environment perception due to its ability to provide more realistic geometric perception and its closer integration with downstream…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Mu Chen , Wenyu Chen , Mingchuan Yang , Yuan Zhang , Tao Han , Xinchi Li , Yunlong Li , Huaici Zhao

3D semantic occupancy prediction is crucial for autonomous driving. While multi-modal fusion improves accuracy over vision-only methods, it typically relies on computationally expensive dense voxel or BEV tensors. We present Gau-Occ, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Chengxin Lv , Yihui Li , Hongyu Yang , YunHong Wang

This article introduces BEVPlace++, a novel, fast, and robust LiDAR global localization method for unmanned ground vehicles. It uses lightweight convolutional neural networks (CNNs) on Bird's Eye View (BEV) image-like representations of…

Robotics · Computer Science 2025-06-26 Lun Luo , Si-Yuan Cao , Xiaorui Li , Jintao Xu , Rui Ai , Zhu Yu , Xieyuanli Chen

3D occupancy, an advanced perception technology for driving scenarios, represents the entire scene without distinguishing between foreground and background by quantifying the physical space into a grid map. The widely adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Jinke Li , Xiao He , Chonghua Zhou , Xiaoqiang Cheng , Yang Wen , Dan Zhang

Integrating LiDAR and Camera information into Bird's-Eye-View (BEV) has become an essential topic for 3D object detection in autonomous driving. Existing methods mostly adopt an independent dual-branch framework to generate LiDAR and camera…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Hongxiang Cai , Zeyuan Zhang , Zhenyu Zhou , Ziyin Li , Wenbo Ding , Jiuhua Zhao

The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Haiming Zhang , Ying Xue , Xu Yan , Jiacheng Zhang , Weichao Qiu , Dongfeng Bai , Bingbing Liu , Shuguang Cui , Zhen Li

This paper addresses the challenge of robotic grasping of general objects. Similar to prior research, the task reads a single-view 3D observation (i.e., point clouds) captured by a depth camera as input. Crucially, the success of object…

Robotics · Computer Science 2024-07-23 Kangqi Ma , Hao Dong , Yadong Mu

We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of free-form language queries. This is a challenging problem…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Antonin Vobecky , Oriane Siméoni , David Hurych , Spyros Gidaris , Andrei Bursuc , Patrick Pérez , Josef Sivic

3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Zijian Zhu , Yichi Zhang , Hai Chen , Yinpeng Dong , Shu Zhao , Wenbo Ding , Jiachen Zhong , Shibao Zheng

Single camera 3D perception for traffic monitoring faces significant challenges due to occlusion and limited field of view. Moreover, fusing information from multiple cameras at the image feature level is difficult because of different view…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Arpitsinh Vaghela , Duo Lu , Aayush Atul Verma , Bharatesh Chakravarthi , Hua Wei , Yezhou Yang

Multi-modal sensor fusion in Bird's Eye View (BEV) representation has become the leading approach for 3D object detection. However, existing methods often rely on depth estimators or transformer encoders to transform image features into BEV…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Yongjin Lee , Hyeon-Mun Jeong , Yurim Jeon , Sanghyun Kim

The application of vision-based multi-view environmental perception system has been increasingly recognized in autonomous driving technology, especially the BEV-based models. Current state-of-the-art solutions primarily encode image…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Di Wu , Feng Yang , Benlian Xu , Pan Liao , Wenhui Zhao , Dingwen Zhang

Bird's Eye View (BEV) is a popular representation for processing 3D point clouds, and by its nature is fundamentally sparse. Motivated by the computational limitations of mobile robot platforms, we create a fast, high-performance BEV 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-08-02 Kyle Vedder , Eric Eaton

Semantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Ruoyu Wang , Yukai Ma , Yi Yao , Sheng Tao , Haoang Li , Zongzhi Zhu , Yong Liu , Xingxing Zuo