English
Related papers

Related papers: OccFormer: Dual-path Transformer for Vision-based …

200 papers

3D semantic occupancy prediction is essential for achieving safe, reliable autonomous driving and robotic navigation. Compared to camera-only perception systems, multi-modal pipelines, especially LiDAR-camera fusion methods, can produce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Lingjun Zhao , Sizhe Wei , James Hays , Lu Gan

Accurate perception of the surrounding environment is essential for safe autonomous driving. 3D occupancy prediction, which estimates detailed 3D structures of roads, buildings, and other objects, is particularly important for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Chihiro Noguchi , Takaki Yamamoto

Vision-based 3D semantic scene completion (SSC) describes autonomous driving scenes through 3D volume representations. However, the occlusion of invisible voxels by scene surfaces poses challenges to current SSC methods in hallucinating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Xiao Zhao , Bo Chen , Mingyang Sun , Dingkang Yang , Youxing Wang , Xukun Zhang , Mingcheng Li , Dongliang Kou , Xiaoyi Wei , Lihua Zhang

3D semantic occupancy prediction offers an intuitive and efficient scene understanding and has attracted significant interest in autonomous driving perception. Existing approaches either rely on full supervision, which demands costly…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Naiyu Fang , Zheyuan Zhou , Fayao Liu , Xulei Yang , Jiacheng Wei , Lemiao Qiu , Hongsheng Li , Guosheng Lin

3D semantic occupancy prediction aims to obtain 3D fine-grained geometry and semantics of the surrounding scene and is an important task for the robustness of vision-centric autonomous driving. Most existing methods employ dense grids such…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yuanhui Huang , Wenzhao Zheng , Yunpeng Zhang , Jie Zhou , Jiwen Lu

3D semantic occupancy prediction is a pivotal task in autonomous driving, providing a dense and fine-grained understanding of the surrounding environment, yet single-modality methods face trade-offs between camera semantics and LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 A. Enes Doruk , Hasan F. Ates

Existing solutions for 3D semantic occupancy prediction typically treat the task as a one-shot 3D voxel-wise segmentation perception problem. These discriminative methods focus on learning the mapping between the inputs and occupancy map in…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Guoqing Wang , Zhongdao Wang , Pin Tang , Jilai Zheng , Xiangxuan Ren , Bailan Feng , Chao Ma

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Dubing Chen , Jin Fang , Wencheng Han , Xinjing Cheng , Junbo Yin , Chenzhong Xu , Fahad Shahbaz Khan , Jianbing Shen

3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods employ dense voxel-based scene representations, ignoring the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Sicheng Zuo , Wenzhao Zheng , Xiaoyong Han , Longchao Yang , Yong Pan , Jiwen Lu

Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects. However, cumbersome voxel features and 3D convolution…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Jinqing Zhang , Yanan Zhang , Qingjie Liu , Yunhong Wang

We present WidthFormer, a novel transformer-based module to compute Bird's-Eye-View (BEV) representations from multi-view cameras for real-time autonomous-driving applications. WidthFormer is computationally efficient, robust and does not…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Chenhongyi Yang , Tianwei Lin , Lichao Huang , Elliot J. Crowley

Occupancy Network has recently attracted much attention in autonomous driving. Instead of monocular 3D detection and recent bird's eye view(BEV) models predicting 3D bounding box of obstacles, Occupancy Network predicts the category of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Mingjie Lu , Yuanxian Huang , Ji Liu , Xingliang Huang , Dong Li , Jinzhang Peng , Lu Tian , Emad Barsoum

Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Yiming Li , Zhiding Yu , Christopher Choy , Chaowei Xiao , Jose M. Alvarez , Sanja Fidler , Chen Feng , Anima Anandkumar

Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key challenges: (1) the…

Artificial Intelligence · Computer Science 2025-09-09 Ruixun Liu , Lingyu Kong , Derun Li , Hang Zhao

Semantic segmentation in autonomous driving has been undergoing an evolution from sparse point segmentation to dense voxel segmentation, where the objective is to predict the semantic occupancy of each voxel in the concerned 3D space. The…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Sicheng Zuo , Wenzhao Zheng , Yuanhui Huang , Jie Zhou , Jiwen Lu

3D environment recognition is essential for autonomous driving systems, as autonomous vehicles require a comprehensive understanding of surrounding scenes. Recently, the predominant approach to define this real-life problem is through 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Huizhou Chen , Jiangyi Wang , Yuxin Li , Na Zhao , Jun Cheng , Xulei Yang

This technical report summarizes the winning solution for the 3D Occupancy Prediction Challenge, which is held in conjunction with the CVPR 2023 Workshop on End-to-End Autonomous Driving and CVPR 23 Workshop on Vision-Centric Autonomous…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Zhiqi Li , Zhiding Yu , David Austin , Mingsheng Fang , Shiyi Lan , Jan Kautz , Jose M. Alvarez

Human driver can easily describe the complex traffic scene by visual system. Such an ability of precise perception is essential for driver's planning. To achieve this, a geometry-aware representation that quantizes the physical 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Chonghao Sima , Wenwen Tong , Tai Wang , Li Chen , Silei Wu , Hanming Deng , Yi Gu , Lewei Lu , Ping Luo , Dahua Lin , Hongyang Li

3D semantic occupancy prediction is a pivotal task in the field of autonomous driving. Recent approaches have made great advances in 3D semantic occupancy predictions on a single modality. However, multi-modal semantic occupancy prediction…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Jingyi Pan , Zipeng Wang , Lin Wang

3D scene understanding plays a vital role in vision-based autonomous driving. While most existing methods focus on 3D object detection, they have difficulty describing real-world objects of arbitrary shapes and infinite classes. Towards a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Yi Wei , Linqing Zhao , Wenzhao Zheng , Zheng Zhu , Jie Zhou , Jiwen Lu