English
Related papers

Related papers: Beyond BEV: Optimizing Point-Level Tokens for Coll…

200 papers

Curating high-quality, domain-specific datasets is a major bottleneck for deploying robust vision systems, requiring complex trade-offs between data quality, diversity, and cost when researching vast, unlabeled data lakes. We introduce…

Bird's-eye-view (BEV) grid is a typical representation of the perception of road components, e.g., drivable area, in autonomous driving. Most existing approaches rely on cameras only to perform segmentation in BEV space, which is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Shubhankar Borse , Senthil Yogamani , Marvin Klingner , Varun Ravi , Hong Cai , Abdulaziz Almuzairee , Fatih Porikli

Pre-training is crucial in 3D-related fields such as autonomous driving where point cloud annotation is costly and challenging. Many recent studies on point cloud pre-training, however, have overlooked the issue of incompleteness, where…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Hao Yang , Haiyang Wang , Di Dai , Liwei Wang

Vision encoders serve as the cornerstone of multimodal understanding. Single-encoder architectures like CLIP exhibit inherent constraints in generalizing across diverse multimodal tasks, while recent multi-encoder fusion methods introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yuchen Liu , Yaoming Wang , Bowen Shi , Xiaopeng Zhang , Wenrui Dai , Chenglin Li , Hongkai Xiong , Qi Tian

Self-attention mechanisms model long-range context by using pairwise attention between all input tokens. In doing so, they assume a fixed attention granularity defined by the individual tokens (e.g., text characters or image pixels), which…

Machine Learning · Computer Science 2022-07-06 Chen Huang , Walter Talbott , Navdeep Jaitly , Josh Susskind

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Rohit Mohan , Daniele Cattaneo , Florian Drews , Abhinav Valada

Collaborative perception is essential for networks of agents with limited sensing capabilities, enabling them to work together by exchanging information to achieve a robust and comprehensive understanding of their environment. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Jiuwu Hao , Liguo Sun , Ti Xiang , Yuting Wan , Haolin Song , Pin Lv

Collaborative perception has garnered significant attention as a crucial technology to overcome the perceptual limitations of single-agent systems. Many state-of-the-art (SOTA) methods have achieved communication efficiency and high…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Gong Chen , Chaokun Zhang , Pengcheng Lv , Xiaohui Xie

3D object detection plays a fundamental role in enabling autonomous driving, which is regarded as the significant key to unlocking the bottleneck of contemporary transportation systems from the perspectives of safety, mobility, and…

Computer Vision and Pattern Recognition · Computer Science 2023-02-08 Zhengwei Bai , Guoyuan Wu , Matthew J. Barth , Yongkang Liu , Emrah Akin Sisbot , Kentaro Oguchi

Vehicle-to-Vehicle (V2V) cooperative perception has great potential to enhance autonomous driving performance by overcoming perception limitations in complex adverse traffic scenarios (CATS). Meanwhile, data serves as the fundamental…

Vehicle-to-Vehicle technologies have enabled autonomous vehicles to share information to see through occlusions, greatly enhancing perception performance. Nevertheless, existing works all focused on homogeneous traffic where vehicles are…

Computer Vision and Pattern Recognition · Computer Science 2023-04-24 Hao Xiang , Runsheng Xu , Jiaqi Ma

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird's-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Jiahui Fu , Chen Gao , Zitian Wang , Lirong Yang , Xiaofei Wang , Beipeng Mu , Si Liu

3D object detection based on LiDAR point clouds is a crucial module in autonomous driving particularly for long range sensing. Most of the research is focused on achieving higher accuracy and these models are not optimized for deployment on…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Sambit Mohapatra , Senthil Yogamani , Heinrich Gotzig , Stefan Milz , Patrick Mader

Efficient relocalization is essential for intelligent vehicles when GPS reception is insufficient or sensor-based localization fails. Recent advances in Bird's-Eye-View (BEV) segmentation allow for accurate estimation of local scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Andrea Boscolo Camiletto , Alfredo Bochicchio , Alexander Liniger , Dengxin Dai , Abel Gawel

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Bohan Li , Yasheng Sun , Zhujin Liang , Dalong Du , Zhuanghui Zhang , Xiaofeng Wang , Yunnan Wang , Xin Jin , Wenjun Zeng

Category-level articulated object pose estimation focuses on the pose estimation of unknown articulated objects within known categories. Despite its significance, this task remains challenging due to the varying shapes and poses of objects,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yuchen Che , Ryo Furukawa , Asako Kanezaki

Accurate positioning is known to be a fundamental requirement for the deployment of Connected Automated Vehicles (CAVs). To meet this need, a new emerging trend is represented by cooperative methods where vehicles fuse information coming…

Signal Processing · Electrical Eng. & Systems 2024-10-28 Luca Barbieri , Bernardo Camajori Tedeschini , Mattia Brambilla , Monica Nicoli

3D perception based on the representations learned from multi-camera bird's-eye-view (BEV) is trending as cameras are cost-effective for mass production in autonomous driving industry. However, there exists a distinct performance gap…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Zeyu Wang , Dingwen Li , Chenxu Luo , Cihang Xie , Xiaodong Yang

Semantic understanding of 3D point clouds is important for various robotics applications. Given that point-wise semantic annotation is expensive, in this paper, we address the challenge of learning models with extremely sparse labels. The…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Liyi Luo , Beiwen Tian , Hao Zhao , Guyue Zhou

Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, often lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jingyu Lei , Gaoang Wang , Der-Horng Lee