English
Related papers

Related papers: FB-OCC: 3D Occupancy Prediction based on Forward-B…

200 papers

End-to-end autonomous driving systems increasingly rely on vision-centric world models to understand and predict their environment. However, a common ineffectiveness in these models is the full reconstruction of future scenes, which expends…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Jianbiao Mei , Yu Yang , Xuemeng Yang , Licheng Wen , Jiajun Lv , Botian Shi , Yong Liu

Multimodal 3D occupancy prediction has garnered significant attention for its potential in autonomous driving. However, most existing approaches are single-modality: camera-based methods lack depth information, while LiDAR-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Zaipeng Duan , Chenxu Dang , Xuzhong Hu , Pei An , Junfeng Ding , Jie Zhan , Yunbiao Xu , Jie Ma

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Dubing Chen , Jin Fang , Wencheng Han , Xinjing Cheng , Junbo Yin , Chenzhong Xu , Fahad Shahbaz Khan , Jianbing Shen

Predicting variations in complex traffic environments is crucial for the safety of autonomous driving. Recent advancements in occupancy forecasting have enabled forecasting future 3D occupied status in driving environments by observing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Junliang Chen , Huaiyuan Xu , Yi Wang , Lap-Pui Chau

3D semantic occupancy prediction is an essential part of autonomous driving, focusing on capturing the geometric details of scenes. Off-road environments are rich in geometric information, therefore it is suitable for 3D semantic occupancy…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Heng Zhai , Jilin Mei , Chen Min , Liang Chen , Fangzhou Zhao , Yu Hu

While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally constrained by environmental variability. Specifically, camera sensors suffer from severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 A. Enes Doruk , Abdelaziz Hussein , Hasan F. Ates

Self-supervision for semantic occupancy estimation is appealing as it removes the labour-intensive manual annotation, thus allowing one to scale to larger autonomous driving datasets. Superquadrics offer an expressive shape family very…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Seamie Hayes , Alexandre Boulch , Andrei Bursuc , Reenu Mohandas , Ganesh Sistu , Tim Brophy , Ciaran Eising

3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subsequent operations like offset learning, attention weighting, and cross-camera aggregation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xun Chen , Tianchen Deng , Rui Wang , Fangjinhua Wang , Junyi Ma , Hongming Shen , Hesheng Wang , Danwei Wang

Autonomous driving without high-definition (HD) maps demands a higher level of active scene understanding. In this competition, the organizers provided the multi-perspective camera images and standard-definition (SD) maps to explore the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Zhongyu Yang , Mai Liu , Jinluo Xie , Yueming Zhang , Chen Shen , Wei Shao , Jichao Jiao , Tengfei Xing , Runbo Hu , Pengfei Xu

Current perception models in autonomous driving heavily rely on large-scale labelled 3D data, which is both costly and time-consuming to annotate. This work proposes a solution to reduce the dependence on labelled 3D training data by…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Chen Min , Xinli Xu , Dawei Zhao , Liang Xiao , Yiming Nie , Bin Dai

Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Zhiwei Lin , Yongtao Wang , Shengxiang Qi , Nan Dong , Ming-Hsuan Yang

Autonomous driving requires a persistent understanding of 3D scenes that is robust to temporal disturbances and accounts for potential future actions. We introduce a new concept of 4D Occupancy Spatio-Temporal Persistence (OccSTeP), which…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Zheng , Jie Hu , Kailun Yang , Jiaming Zhang

Seeing only a tiny part of the whole is not knowing the full circumstance. Bird's-eye-view (BEV) perception, a process of obtaining allocentric maps from egocentric views, is restricted when using a narrow Field of View (FoV) alone. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Zhifeng Teng , Jiaming Zhang , Kailun Yang , Kunyu Peng , Hao Shi , Simon Reiß , Ke Cao , Rainer Stiefelhagen

Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to facilitate the learning, but often fail to fully exploit the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Liang Peng , Junkai Xu , Haoran Cheng , Zheng Yang , Xiaopei Wu , Wei Qian , Wenxiao Wang , Boxi Wu , Deng Cai

Developing 3D semantic occupancy prediction models often relies on dense 3D annotations for supervised learning, a process that is both labor and resource-intensive, underscoring the need for label-efficient or even label-free approaches.…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Samuel Sze , Daniele De Martini , Lars Kunze

Bird's-eye-view (BEV) semantic segmentation is becoming crucial in autonomous driving systems. It realizes ego-vehicle surrounding environment perception by projecting 2D multi-view images into 3D world space. Recently, BEV segmentation has…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Jian Sun , Yuqi Dai , Chi-Man Vong , Qing Xu , Shengbo Eben Li , Jianqiang Wang , Lei He , Keqiang Li

The advent of deep learning has led to significant progress in monocular human reconstruction. However, existing representations, such as parametric models, voxel grids, meshes and implicit neural representations, have difficulties…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Qiao Feng , Yebin Liu , Yu-Kun Lai , Jingyu Yang , Kun Li

Estimating 3D occupancy and motion at the vehicle's surroundings is essential for autonomous driving, enabling situational awareness in dynamic environments. Existing approaches jointly learn geometry and motion but rely on expensive 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xavier Timoneda , Markus Herb , Fabian Duerr , Daniel Goehring

Detection of moving objects is a very important task in autonomous driving systems. After the perception phase, motion planning is typically performed in Bird's Eye View (BEV) space. This would require projection of objects detected on the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Hazem Rashed , Mariam Essam , Maha Mohamed , Ahmad El Sallab , Senthil Yogamani

Language-conditioned local navigation requires a robot to infer a nearby traversable target location from its current observation and an open-vocabulary, relational instruction. Existing vision-language spatial grounding methods usually…

Robotics · Computer Science 2026-03-11 Xinyu Gao , Gang Chen , Javier Alonso-Mora