English
Related papers

Related papers: LaRa: Latents and Rays for Multi-Camera Bird's-Eye…

200 papers

Bird's-Eye-View (BEV) maps have emerged as one of the most powerful representations for scene understanding due to their ability to provide rich spatial context while being easy to interpret and process. Such maps have found use in many…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Nikhil Gosala , Abhinav Valada

Bird's-Eye-View (BEV) semantic segmentation provides comprehensive environmental perception for autonomous driving but suffers multi-modal misalignment and sensor noise. We propose RESAR-BEV, a progressive refinement framework that advances…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Zhiwen Zeng , Yunfei Yin , Zheng Yuan , Argho Dey , Xianjian Bao

Modern methods for vision-centric autonomous driving perception widely adopt the bird's-eye-view (BEV) representation to describe a 3D scene. Despite its better efficiency than voxel representation, it has difficulty describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Yuanhui Huang , Wenzhao Zheng , Yunpeng Zhang , Jie Zhou , Jiwen Lu

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Zhiqi Li , Wenhai Wang , Hongyang Li , Enze Xie , Chonghao Sima , Tong Lu , Qiao Yu , Jifeng Dai

Referring image segmentation aims to produce a pixel-level mask for the image region described by a natural-language expression. Although pretrained vision-language models have improved semantic grounding, many existing methods still rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Alaa Dalaq , Muzammil Behzad

Bird's eye view (BEV) perception is becoming increasingly important in the field of autonomous driving. It uses multi-view camera data to learn a transformer model that directly projects the perception of the road environment onto the BEV…

Computer Vision and Pattern Recognition · Computer Science 2023-09-11 Rui Song , Runsheng Xu , Andreas Festag , Jiaqi Ma , Alois Knoll

Moving object detection and segmentation is an essential task in the Autonomous Driving pipeline. Detecting and isolating static and moving components of a vehicle's surroundings are particularly crucial in path planning and localization…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Sambit Mohapatra , Mona Hodaei , Senthil Yogamani , Stefan Milz , Heinrich Gotzig , Martin Simon , Hazem Rashed , Patrick Maeder

Birds' Eye View (BEV) semantic segmentation is an indispensable perception task in end-to-end autonomous driving systems. Unsupervised and semi-supervised learning for BEV tasks, as pivotal for real-world applications, underperform due to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Siyu Li , Fei Teng , Yihong Cao , Kailun Yang , Zhiyong Li , Yaonan Wang

Accurate prediction of communication link quality metrics is essential for vehicle-to-infrastructure (V2I) systems, enabling smooth handovers, efficient beam management, and reliable low-latency communication. The increasing availability of…

Machine Learning · Computer Science 2025-09-05 Kimia Ehsani , Walid Saad

Road intersection monitoring and control research often utilize bird's eye view (BEV) simulators. In real traffic settings, achieving a BEV akin to that in a simulator necessitates the deployment of drones or specific sensor mounting, which…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Rukesh Prajapati , Amr S. El-Wakeel

Bird's-eye-view (BEV) grid is a common representation for the perception of road components, e.g., drivable area, in autonomous driving. Most existing approaches rely on cameras only to perform segmentation in BEV space, which is…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Shubhankar Borse , Marvin Klingner , Varun Ravi Kumar , Hong Cai , Abdulaziz Almuzairee , Senthil Yogamani , Fatih Porikli

Trajectory prediction is, naturally, a key task for vehicle autonomy. While the number of traffic rules is limited, the combinations and uncertainties associated with each agent's behaviour in real-world scenarios are nearly impossible to…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Sushil Sharma , Arindam Das , Ganesh Sistu , Mark Halton , Ciarán Eising

This paper proposes an efficient multi-camera to Bird's-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from poor transformation efficiency or rely on device-specific…

Computer Vision and Pattern Recognition · Computer Science 2022-11-22 Hongyu Zhou , Zheng Ge , Zeming Li , Xiangyu Zhang

Bird's-eye-view (BEV) grid is a typical representation of the perception of road components, e.g., drivable area, in autonomous driving. Most existing approaches rely on cameras only to perform segmentation in BEV space, which is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Shubhankar Borse , Senthil Yogamani , Marvin Klingner , Varun Ravi , Hong Cai , Abdulaziz Almuzairee , Fatih Porikli

Understanding the scene around the ego-vehicle is key to assisted and autonomous driving. Nowadays, this is mostly conducted using cameras and laser scanners, despite their reduced performances in adverse weather conditions. Automotive…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Arthur Ouaknine , Alasdair Newson , Patrick Pérez , Florence Tupin , Julien Rebut

LiDAR and camera are two essential sensors for 3D object detection in autonomous driving. LiDAR provides accurate and reliable 3D geometry information while the camera provides rich texture with color. Despite the increasing popularity of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Qi Jiang , Hao Sun , Xi Zhang

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xiaodong Mei , Diankun Zhang , Hongwei Xie , Guang Chen , Hangjun Ye , Dan Xu

It is a crucial step to achieve effective semantic segmentation of lane marking during the construction of the lane level high-precision map. In recent years, many image semantic segmentation methods have been proposed. These methods mainly…

Computer Vision and Pattern Recognition · Computer Science 2020-03-11 Ruochen Yin , Biao Yu , Huapeng Wu , Yutao Song , Runxin Niu

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We…

3D object detection from LiDAR data for autonomous driving has been making remarkable strides in recent years. Among the state-of-the-art methodologies, encoding point clouds into a bird's eye view (BEV) has been demonstrated to be both…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Yantao Lu , Xuetao Hao , Yilan Li , Weiheng Chai , Shiqi Sun , Senem Velipasalar