English
Related papers

Related papers: RoboFusion: Towards Robust Multi-Modal 3D Object D…

200 papers

To reduce the amount of transmitted data, feature map based fusion is recently proposed as a practical solution to cooperative 3D object detection by autonomous vehicles. The precision of object detection, however, may require significant…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Jingda Guo , Dominic Carrillo , Sihai Tang , Qi Chen , Qing Yang , Song Fu , Xi Wang , Nannan Wang , Paparao Palacharla

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically back-project and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tomas Berriel Martins , Martin R. Oswald , Javier Civera

Unmanned aerial vehicle (UAV) object detection plays a vital role in applications such as environmental monitoring and urban security. To improve robustness, recent studies have explored multimodal detection by fusing visible (RGB) and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Liu Zongzhen , Luo Hui , Wang Zhixing , Wei Yuxing , Zuo Haorui , Zhang Jianlin

Current LiDAR-only 3D detection methods inevitably suffer from the sparsity of point clouds. Many multi-modal methods are proposed to alleviate this issue, while different representations of images and point clouds make it difficult to fuse…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Xiaopei Wu , Liang Peng , Honghui Yang , Liang Xie , Chenxi Huang , Chengqi Deng , Haifeng Liu , Deng Cai

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion…

Safety of the Intended Functionality (SOTIF) addresses sensor performance limitations and deep learning-based object detection insufficiencies to ensure the intended functionality of Automated Driving Systems (ADS). This paper presents a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Milin Patel , Rolf Jung

Environmental perception with the multi-modal fusion of radar and camera is crucial in autonomous driving to increase accuracy, completeness, and robustness. This paper focuses on utilizing millimeter-wave (MMW) radar and camera sensor…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Taohua Zhou , Yining Shi , Junjie Chen , Kun Jiang , Mengmeng Yang , Diange Yang

In the recent literature, on the one hand, many 3D multi-object tracking (MOT) works have focused on tracking accuracy and neglected computation speed, commonly by designing rather complex cost functions and feature extractors. On the other…

Computer Vision and Pattern Recognition · Computer Science 2022-08-29 Xiyang Wang , Chunyun Fu , Zhankun Li , Ying Lai , Jiawei He

Existing LiDAR-Camera fusion methods have achieved strong results in 3D object detection. To address the sparsity of point clouds, previous approaches typically construct spatial pseudo point clouds via depth completion as auxiliary input…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Jijun Wang , Yan Wu , Yujian Mo , Junqiao Zhao , Jun Yan , Yinghao Hu

The safety of an automated vehicle hinges crucially upon the accuracy of perception and decision-making latency. Under these stringent requirements, future automated cars are usually equipped with multi-modal sensors such as cameras and…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-15 Zhendong Wang , Xiaoming Zeng , Shuaiwen Leon Song , Yang Hu

This paper presents the Autonomous Driving Segment Anything Model (AD-SAM), a fine-tuned vision foundation model for semantic segmentation in autonomous driving (AD). AD-SAM extends the Segment Anything Model (SAM) with a dual-encoder and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Mario Camarena , Het Patel , Fatemeh Nazari , Evangelos Papalexakis , Mohamadhossein Noruzoliaee , Jia Chen

In autonomous driving, recent research has increasingly focused on collaborative perception based on deep learning to overcome the limitations of individual perception systems. Although these methods achieve high accuracy, they rely on high…

Robotics · Computer Science 2025-07-04 Maryem Fadili , Mohamed Anis Ghaoui , Louis Lecrosnier , Steve Pechberti , Redouane Khemmar

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and…

Artificial Intelligence · Computer Science 2025-09-23 Jia Li , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse…

Computer Vision and Pattern Recognition · Computer Science 2022-10-19 Xin Li , Botian Shi , Yuenan Hou , Xingjiao Wu , Tianlong Ma , Yikang Li , Liang He

4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fuse 4D Radar and camera data confront several challenges. First, their absolute depth…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhongyu Xia , Yousen Tang , Yongtao Wang , Zhifeng Wang , Weijun Qin

Autonomous driving has now made great strides thanks to artificial intelligence, and numerous advanced methods have been proposed for vehicle end target detection, including single sensor or multi sensor detection methods. However, the…

Robotics · Computer Science 2023-06-13 Xiuyu Yang , Zhuangyan Zhang , Haikuo Du , Sui Yang , Fengping Sun , Yanbo Liu , Ling Pei , Wenchao Xu , Weiqi Sun , Zhengyu Li

A significant challenge in object detection is accurate identification of an object's position in image space, whereas one algorithm with one set of parameters is usually not enough, and the fusion of multiple algorithms and/or parameters…

Computer Vision and Pattern Recognition · Computer Science 2018-03-20 Pan Wei , John E. Ball , Derek T. Anderson

In autonomous driving, Vehicle-Infrastructure Cooperative 3D Object Detection (VIC3D) makes use of multi-view cameras from both vehicles and traffic infrastructure, providing a global vantage point with rich semantic context of road…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Zhe Wang , Siqi Fan , Xiaoliang Huo , Tongda Xu , Yan Wang , Jingjing Liu , Yilun Chen , Ya-Qin Zhang

Most autonomous vehicles are equipped with LiDAR sensors and stereo cameras. The former is very accurate but generates sparse data, whereas the latter is dense, has rich texture and color information but difficult to extract robust 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-11-10 Farzin Negahbani , Onur Berk Töre , Fatma Güney , Baris Akgun

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei
‹ Prev 1 4 5 6 7 8 10 Next ›