中文
相关论文

相关论文: Towards Domain Generalization for Multi-view 3D Ob…

200 篇论文

3D visual perception tasks based on multi-camera images are essential for autonomous driving systems. Latest work in this field performs 3D object detection by leveraging multi-view images as an input and iteratively enhancing object…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Jongwoo Park , Apoorv Singh , Varun Bankiti

Supervised 3D Object Detection models have been displaying increasingly better performance in single-domain cases where the training data comes from the same environment and sensor as the testing data. However, in real-world scenarios data…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Louis Soum-Fontez , Jean-Emmanuel Deschaud , François Goulette

In this paper, we present BEVerse, a unified framework for 3D perception and prediction based on multi-camera systems. Unlike existing studies focusing on the improvement of single-task approaches, BEVerse features in producing…

计算机视觉与模式识别 · 计算机科学 2022-05-20 Yunpeng Zhang , Zheng Zhu , Wenzhao Zheng , Junjie Huang , Guan Huang , Jie Zhou , Jiwen Lu

Modern autonomous driving systems increasingly rely on mixed camera configurations with pinhole and fisheye cameras for full view perception. However, Bird's-Eye View (BEV) 3D object detection models are predominantly designed for pinhole…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Xiangzhong Liu , Hao Shen

Deep learning models such as convolutional neural networks and transformers have been widely applied to solve 3D object detection problems in the domain of autonomous driving. While existing models have achieved outstanding performance on…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Ruixiao Zhang , Juheon Lee , Xiaohao Cai , Adam Prugel-Bennett

Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detection and segmentation features are geometrically registered to the same physical…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Ahmet İnanç , Özgür Erkent

In this paper, we propose a novel network framework for indoor 3D object detection to handle variable input frame numbers in practical scenarios. Existing methods only consider fixed frames of input data for a single detector, such as…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Zhenyu Wu , Xiuwei Xu , Ziwei Wang , Chong Xia , Linqing Zhao , Jiwen Lu , Haibin Yan

Learning-based monocular depth estimation leverages geometric priors present in the training data to enable metric depth perception from a single image, a traditionally ill-posed problem. However, these priors are often specific to a…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Karlo Koledić , Luka Petrović , Ivan Petrović , Ivan Marković

In current research, Bird's-Eye-View (BEV)-based transformers are increasingly utilized for multi-camera 3D object detection. Traditional models often employ random queries as anchors, optimizing them successively. Recent advancements…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Marius Dähling , Sebastian Krebs , J. Marius Zöllner

Multi-view aggregation promises to overcome the occlusion and missed detection challenge in multi-object detection and tracking. Recent approaches in multi-view detection and 3D object detection made a huge performance leap by projecting…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Torben Teepe , Philipp Wolters , Johannes Gilg , Fabian Herzog , Gerhard Rigoll

Despite the significant improvement in the performance of monocular pose estimation approaches and their ability to generalize to unseen environments, multi-view (MV) approaches are often lagging behind in terms of accuracy and are specific…

计算机视觉与模式识别 · 计算机科学 2019-10-09 Abdolrahim Kadkhodamohammadi , Nicolas Padoy

Monocular 3D object detection (Mono3D) has achieved unprecedented success with the advent of deep learning techniques and emerging large-scale autonomous driving datasets. However, drastic performance degradation remains an unwell-studied…

计算机视觉与模式识别 · 计算机科学 2022-04-28 Zhenyu Li , Zehui Chen , Ang Li , Liangji Fang , Qinhong Jiang , Xianming Liu , Junjun Jiang

Monocular 3D lane detection is challenging due to the difficulty in capturing depth information from single-camera images. A common strategy involves transforming front-view (FV) images into bird's-eye-view (BEV) space through inverse…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Dongxin Lyu , Han Huang , Cheng Tan , Zimu Li

Single Domain Generalization (SDG) tackles the problem of training a model on a single source domain so that it generalizes to any unseen target domain. While this has been well studied for image classification, the literature on SDG object…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Vidit Vidit , Martin Engilberge , Mathieu Salzmann

Camera-based 3D object detection in BEV (Bird's Eye View) space has drawn great attention over the past few years. Dense detectors typically follow a two-stage pipeline by first constructing a dense BEV feature and then performing object…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Haisong Liu , Yao Teng , Tao Lu , Haiguang Wang , Limin Wang

Multi-view camera-based 3D perception can be conducted using bird's eye view (BEV) features obtained through perspective view-to-BEV transformations. Several studies have shown that the performance of these 3D perception methods can be…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Junho Koh , Youngwoo Lee , Jungho Kim , Dongyoung Lee , Jun Won Choi

In this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Changcai Li , Wenwei Lin , Zuoxun Hou , Gang Chen , Wei Zhang , Huihui Zhou , Weishi Zheng

With the attention gained by camera-only 3D object detection in autonomous driving, methods based on Bird-Eye-View (BEV) representation especially derived from the forward view transformation paradigm, i.e., lift-splat-shoot (LSS), have…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Weijie Ma , Jingwei Jiang , Yang Yang , Zehui Chen , Hao Chen

In this paper, we present DAT, a Depth-Aware Transformer framework designed for camera-based 3D detection. Our model is based on observing two major issues in existing methods: large depth translation errors and duplicate predictions along…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Hao Zhang , Hongyang Li , Ailing Zeng , Feng Li , Shilong Liu , Xingyu Liao , Lei Zhang

Monocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image. In this paper, we strive to boost currently underperforming monocular…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Weijia Zhang , Dongnan Liu , Chao Ma , Weidong Cai