中文
相关论文

相关论文: Addressing Diverging Training Costs using BEVResto…

200 篇论文

Modern autonomous driving systems increasingly rely on mixed camera configurations with pinhole and fisheye cameras for full view perception. However, Bird's-Eye View (BEV) 3D object detection models are predominantly designed for pinhole…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Xiangzhong Liu , Hao Shen

Camera-based Bird's-Eye-View (BEV) perception often struggles between adopting 3D-to-2D or 2D-to-3D view transformation (VT). The 3D-to-2D VT typically employs resource-intensive Transformer to establish robust correspondences between 3D…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Peidong Li , Wancheng Shen , Qihao Huang , Dixiao Cui

Bird's eye view (BEV) perception is becoming increasingly important in the field of autonomous driving. It uses multi-view camera data to learn a transformer model that directly projects the perception of the road environment onto the BEV…

计算机视觉与模式识别 · 计算机科学 2023-09-11 Rui Song , Runsheng Xu , Andreas Festag , Jiaqi Ma , Alois Knoll

End-to-end reinforcement learning on images showed significant progress in the recent years. Data-based approach leverage data augmentation and domain randomization while representation learning methods use auxiliary losses to learn…

机器学习 · 计算机科学 2024-01-19 Tom Dupuis , Jaonary Rabarisoa , Quoc-Cuong Pham , David Filliat

Pairwise point cloud registration is a critical task for many applications, which heavily depends on finding correct correspondences from the two point clouds. However, the low overlap between input point clouds causes the registration to…

计算机视觉与模式识别 · 计算机科学 2023-03-15 Lin Li , Wendong Ding , Yongkun Wen , Yufei Liang , Yong Liu , Guowei Wan

Understanding road geometry is a critical component of the autonomous vehicle (AV) stack. While high-definition (HD) maps can readily provide such information, they suffer from high labeling and maintenance costs. Accordingly, many recent…

机器人学 · 计算机科学 2024-07-10 Xunjiang Gu , Guanyu Song , Igor Gilitschenski , Marco Pavone , Boris Ivanovic

Accurate layout estimation is crucial for planning and navigation in robotics applications, such as self-driving. In this paper, we introduce the Stereo Bird's Eye ViewNetwork (SBEVNet), a novel supervised end-to-end framework for…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Divam Gupta , Wei Pu , Trenton Tabor , Jeff Schneider

Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Zhiwei Lin , Yongtao Wang , Shengxiang Qi , Nan Dong , Ming-Hsuan Yang

This paper introduces BEV-VLM, a novel approach for trajectory planning in autonomous driving that leverages Vision-Language Models (VLMs) with Bird's-Eye View (BEV) feature maps as visual input. Unlike conventional trajectory planning…

机器人学 · 计算机科学 2026-03-02 Guancheng Chen , Sheng Yang , Tong Zhan , Jian Wang

3D object detection plays a pivotal role in autonomous driving and robotics, demanding precise interpretation of Bird's Eye View (BEV) images. The dynamic nature of real-world environments necessitates the use of dynamic query mechanisms in…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Jiawei Yao , Yingxin Lai , Hongrui Kou , Tong Wu , Ruixi Liu

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond…

计算机视觉与模式识别 · 计算机科学 2023-04-12 Lei Yang , Kaicheng Yu , Tao Tang , Jun Li , Kun Yuan , Li Wang , Xinyu Zhang , Peng Chen

Autonomous driving requires an accurate representation of the environment. A strategy toward high accuracy is to fuse data from several sensors. Learned Bird's-Eye View (BEV) encoders can achieve this by mapping data from individual sensors…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Thomas Monninger , Vandana Dokkadi , Md Zafar Anwar , Steffen Staab

Bird's-Eye-View (BEV) 3D Object Detection is a crucial multi-view technique for autonomous driving systems. Recently, plenty of works are proposed, following a similar paradigm consisting of three essential components, i.e., camera feature…

计算机视觉与模式识别 · 计算机科学 2022-12-05 Xiaowei Chi , Jiaming Liu , Ming Lu , Rongyu Zhang , Zhaoqing Wang , Yandong Guo , Shanghang Zhang

In this paper, we present BEVerse, a unified framework for 3D perception and prediction based on multi-camera systems. Unlike existing studies focusing on the improvement of single-task approaches, BEVerse features in producing…

计算机视觉与模式识别 · 计算机科学 2022-05-20 Yunpeng Zhang , Zheng Zhu , Wenzhao Zheng , Junjie Huang , Guan Huang , Jie Zhou , Jiwen Lu

Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Zhiwei Lin , Zhe Liu , Zhongyu Xia , Xinhao Wang , Yongtao Wang , Shengxiang Qi , Yang Dong , Nan Dong , Le Zhang , Ce Zhu

Monocular Visual Odometry (MVO) provides a cost-effective, real-time positioning solution for autonomous vehicles. However, MVO systems face the common issue of lacking inherent scale information from monocular cameras. Traditional methods…

机器人学 · 计算机科学 2025-02-28 Yufei Wei , Sha Lu , Wangtao Lu , Rong Xiong , Yue Wang

Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune…

Bird's eye view (BEV) is widely adopted by most of the current point cloud detectors due to the applicability of well-explored 2D detection techniques. However, existing methods obtain BEV features by simply collapsing voxel or point…

计算机视觉与模式识别 · 计算机科学 2023-01-30 Dihe Huang , Ying Chen , Yikang Ding , Jinli Liao , Jianlin Liu , Kai Wu , Qiang Nie , Yong Liu , Chengjie Wang , Zhiheng Li

Image-to-point cloud cross-modal Visual Place Recognition (VPR) is a challenging task where the query is an RGB image, and the database samples are LiDAR point clouds. Compared to single-modal VPR, this approach benefits from the widespread…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Jianyi Peng , Fan Lu , Bin Li , Yuan Huang , Sanqing Qu , Guang Chen

Autonomous driving requires understanding infrastructure elements, such as lanes and crosswalks. To navigate safely, this understanding must be derived from sensor data in real-time and needs to be represented in vectorized form. Learned…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Thomas Monninger , Md Zafar Anwar , Stanislaw Antol , Steffen Staab , Sihao Ding