中文
相关论文

相关论文: OnlineBEV: Recurrent Temporal Fusion in Bird's Eye…

200 篇论文

Autonomous vehicles (AV) require that neural networks used for perception be robust to different viewpoints if they are to be deployed across many types of vehicles without the repeated cost of data collection and labeling for each. AV…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Tzofi Klinghoffer , Jonah Philion , Wenzheng Chen , Or Litany , Zan Gojcic , Jungseock Joo , Ramesh Raskar , Sanja Fidler , Jose M. Alvarez

Bounded by the inherent ambiguity of depth perception, contemporary camera-based 3D object detection methods fall into the performance bottleneck. Intuitively, leveraging temporal multi-view stereo (MVS) technology is the natural knowledge…

计算机视觉与模式识别 · 计算机科学 2022-09-22 Yinhao Li , Han Bao , Zheng Ge , Jinrong Yang , Jianjian Sun , Zeming Li

Autonomous driving perceives its surroundings for decision making, which is one of the most complex scenarios in visual perception. The success of paradigm innovation in solving the 2D object detection task inspires us to seek an elegant,…

计算机视觉与模式识别 · 计算机科学 2022-06-17 Junjie Huang , Guan Huang , Zheng Zhu , Yun Ye , Dalong Du

3D object detection based on LiDAR point clouds is a crucial module in autonomous driving particularly for long range sensing. Most of the research is focused on achieving higher accuracy and these models are not optimized for deployment on…

计算机视觉与模式识别 · 计算机科学 2021-07-13 Sambit Mohapatra , Senthil Yogamani , Heinrich Gotzig , Stefan Milz , Patrick Mader

Most automated driving systems comprise a diverse sensor set, including several cameras, Radars, and LiDARs, ensuring a complete 360\deg coverage in near and far regions. Unlike Radar and LiDAR, which measure directly in 3D, cameras capture…

机器人学 · 计算机科学 2023-09-20 David Unger , Nikhil Gosala , Varun Ravi Kumar , Shubhankar Borse , Abhinav Valada , Senthil Yogamani

Recent 3D object detectors typically utilize multi-sensor data and unify multi-modal features in the shared bird's-eye view (BEV) representation space. However, our empirical findings indicate that previous methods have limitations in…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Jiahui Fu , Chen Gao , Zitian Wang , Lirong Yang , Xiaofei Wang , Beipeng Mu , Si Liu

Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods transform multimodal data from independent sensors, such as…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Markus Essl , Marta Moscati , Mubashir Noman , Muhammad Zaigham Zaheer , Usman Naseem , Shah Nawaz , Markus Schedl

As a cornerstone technique for autonomous driving, Bird's Eye View (BEV) segmentation has recently achieved remarkable progress with pinhole cameras. However, it is non-trivial to extend the existing methods to fisheye cameras with severe…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Hang Li , Dianmo Sheng , Qiankun Dong , Zichun Wang , Zhiwei Xu , Tao Li

In autonomous driving, accurate 3D lane detection using monocular cameras is important for downstream tasks. Recent CNN and Transformer approaches usually apply a two-stage model design. The first stage transforms the image feature from a…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yifeng Bai , Zhirong Chen , Pengpeng Liang , Bo Song , Erkang Cheng

In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being…

计算机视觉与模式识别 · 计算机科学 2023-06-08 Yuexin Ma , Tai Wang , Xuyang Bai , Huitong Yang , Yuenan Hou , Yaming Wang , Yu Qiao , Ruigang Yang , Dinesh Manocha , Xinge Zhu

In spite of the recent advancements in multi-object tracking, occlusion poses a significant challenge. Multi-camera setups have been used to address this challenge by providing a comprehensive coverage of the scene. Recent multi-view…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Reef Alturki , Adrian Hilton , Jean-Yves Guillemaut

While recent camera-only 3D detection methods leverage multiple timesteps, the limited history they use significantly hampers the extent to which temporal fusion can improve object perception. Observing that existing works' fusion of…

计算机视觉与模式识别 · 计算机科学 2022-10-06 Jinhyung Park , Chenfeng Xu , Shijia Yang , Kurt Keutzer , Kris Kitani , Masayoshi Tomizuka , Wei Zhan

Accurate, fast, and reliable 3D perception is essential for autonomous driving. Recently, bird's-eye view (BEV)-based perception approaches have emerged as superior alternatives to perspective-based solutions, offering enhanced spatial…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Ozsel Kilinc , Cem Tarhan

A robust awareness of how dynamic scenes evolve is essential for Autonomous Driving systems, as they must accurately detect, track, and predict the behaviour of surrounding obstacles. Traditional perception pipelines that rely on modular…

In this paper, we propose PETRv2, a unified framework for 3D perception from multi-view images. Based on PETR, PETRv2 explores the effectiveness of temporal modeling, which utilizes the temporal information of previous frames to boost 3D…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Yingfei Liu , Junjie Yan , Fan Jia , Shuailin Li , Aqi Gao , Tiancai Wang , Xiangyu Zhang , Jian Sun

Existing approaches to drone visual geo-localization predominantly adopt the image-based setting, where a single drone-view snapshot is matched with images from other platforms. Such task formulation, however, underutilizes the inherent…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Hao Ju , Shaofei Huang , Si Liu , Zhedong Zheng

Camera-only Bird's Eye View (BEV) has demonstrated great potential in environment perception in a 3D space. However, most existing studies were conducted under a supervised setup which cannot scale well while handling various new data.…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Kai Jiang , Jiaxing Huang , Weiying Xie , Yunsong Li , Ling Shao , Shijian Lu

Transforming image features from perspective view (PV) space to bird's-eye-view (BEV) space remains challenging in autonomous driving due to depth ambiguity and occlusion. Although several view transformation (VT) paradigms have been…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Jeongbin Hong , Dooseop Choi , Taeg-Hyun An , Kyounghwan An , Kyoung-Wook Min

Recent deep learning models achieve impressive results on 3D scene analysis tasks by operating directly on unstructured point clouds. A lot of progress was made in the field of object classification and semantic segmentation. However, the…

计算机视觉与模式识别 · 计算机科学 2019-12-20 Cathrin Elich , Francis Engelmann , Theodora Kontogianni , Bastian Leibe

A popular approach for constructing bird's-eye-view (BEV) representation in 3D detection is to lift 2D image features onto the viewing frustum space based on explicitly predicted depth distribution. However, depth distribution can only…

计算机视觉与模式识别 · 计算机科学 2024-01-12 Zaibin Zhang , Yuanhang Zhang , Lijun Wang , Yifan Wang , Huchuan Lu