English
Related papers

Related papers: Reconstruction Matters: Learning Geometry-Aligned …

200 papers

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Zhiqi Li , Wenhai Wang , Hongyang Li , Enze Xie , Chonghao Sima , Tong Lu , Qiao Yu , Jifeng Dai

Bird's-Eye View (BEV) features are popular intermediate scene representations shared by the 3D backbone and the detector head in LiDAR-based object detectors. However, little research has been done to investigate how to incorporate…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Haitao Yang , Zaiwei Zhang , Xiangru Huang , Min Bai , Chen Song , Bo Sun , Li Erran Li , Qixing Huang

Self-supervised learning (SSL) for point cloud pre-training has become a cornerstone for many 3D vision tasks, enabling effective learning from large-scale unannotated data. At the scene level, existing SSL methods often incorporate volume…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Keyi Liu , Weidong Yang , Ben Fei , Ying He

Recently, perception task based on Bird's-Eye View (BEV) representation has drawn more and more attention, and BEV representation is promising as the foundation for next-generation Autonomous Vehicle (AV) perception. However, most existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yangguang Li , Bin Huang , Zeren Chen , Yufeng Cui , Feng Liang , Mingzhu Shen , Fenggang Liu , Enze Xie , Lu Sheng , Wanli Ouyang , Jing Shao

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while elevating novel view quality. Due to the surround-view with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Junhong Lin , Kangli Wang , Shunzhou Wang , Songlin Fan , Ge Li , Wei Gao

Learning Bird's Eye View (BEV) representation from surrounding-view cameras is of great importance for autonomous driving. In this work, we propose a Geometry-guided Kernel Transformer (GKT), a novel 2D-to-BEV representation learning…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Shaoyu Chen , Tianheng Cheng , Xinggang Wang , Wenming Meng , Qian Zhang , Wenyu Liu

Bird's-eye-view (BEV) representation is crucial for the perception function in autonomous driving tasks. It is difficult to balance the accuracy, efficiency and range of BEV representation. The existing works are restricted to a limited…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Hang Wu , Zhenghao Zhang , Siyuan Lin , Tong Qin , Jin Pan , Qiang Zhao , Chunjing Xu , Ming Yang

Bird's-Eye-View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature…

Computer Vision and Pattern Recognition · Computer Science 2022-12-05 Xiaowei Chi , Jiaming Liu , Ming Lu , Rongyu Zhang , Zhaoqing Wang , Yandong Guo , Shanghang Zhang

Camera-based Bird's-Eye-View (BEV) perception often struggles between adopting 3D-to-2D or 2D-to-3D view transformation (VT). The 3D-to-2D VT typically employs resource-intensive Transformer to establish robust correspondences between 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-09-16 Peidong Li , Wancheng Shen , Qihao Huang , Dixiao Cui

Bird's-Eye-View (BEV) perception has become a vital component of autonomous driving systems due to its ability to integrate multiple sensor inputs into a unified representation, enhancing performance in various downstream tasks. However,…

Robotics · Computer Science 2024-10-10 Yuxin Li , Yiheng Li , Xulei Yang , Mengying Yu , Zihang Huang , Xiaojun Wu , Chai Kiat Yeo

We investigate data augmentation for 3D object detection in autonomous driving. We utilize recent advancements in 3D reconstruction based on Gaussian Splatting for 3D object placement in driving scenes. Unlike existing diffusion-based…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Farhad G. Zanjani , Davide Abati , Auke Wiggers , Dimitris Kalatzis , Jens Petersen , Hong Cai , Amirhossein Habibian

Bird's-eye-view (BEV) semantic segmentation is becoming crucial in autonomous driving systems. It realizes ego-vehicle surrounding environment perception by projecting 2D multi-view images into 3D world space. Recently, BEV segmentation has…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Jian Sun , Yuqi Dai , Chi-Man Vong , Qing Xu , Shengbo Eben Li , Jianqiang Wang , Lei He , Keqiang Li

3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Zijian Zhu , Yichi Zhang , Hai Chen , Yinpeng Dong , Shu Zhao , Wenbo Ding , Jiachen Zhong , Shibao Zheng

Generating a detailed near-field perceptual model of the environment is an important and challenging problem in both self-driving vehicles and autonomous mobile robotics. A Bird Eye View (BEV) map, providing a panoptic representation, is a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Pramit Dutta , Ganesh Sistu , Senthil Yogamani , Edgar Galván , John McDonald

Visual bird's eye view (BEV) semantic segmentation helps autonomous vehicles understand the surrounding environment only from images, including static elements (e.g., roads) and dynamic elements (e.g., vehicles, pedestrians). However, the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Junyu Zhu , Lina Liu , Yu Tang , Feng Wen , Wanlong Li , Yong Liu

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Bohan Li , Yasheng Sun , Zhujin Liang , Dalong Du , Zhuanghui Zhang , Xiaofeng Wang , Yunnan Wang , Xin Jin , Wenjun Zeng

3D Gaussian Splatting (3DGS) is a recent approach for scene rendering. Although primarily designed for view synthesis, its potential for scene understanding tasks remains underexplored. In this work, we conduct a comparative evaluation of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Julia Farganus , Krzysztof Żurawicki , Arkadiusz Gaweł , Weronika Jakubowska , Halina Kwaśnicka

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Thomas Monninger , Shaoyuan Xie , Qi Alfred Chen , Sihao Ding

Multi-sensor fusion is essential for an accurate and reliable autonomous driving system. Recent approaches are based on point-level fusion: augmenting the LiDAR point cloud with camera features. However, the camera-to-LiDAR projection…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Zhijian Liu , Haotian Tang , Alexander Amini , Xinyu Yang , Huizi Mao , Daniela Rus , Song Han

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. However, we argue…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Zeming Chen , Hang Zhao