English
Related papers

Related papers: Multi-Modal Fusion Transformer for End-to-End Auto…

200 papers

Localization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this…

Robotics · Computer Science 2024-04-08 Chenyang Wu , Yifan Duan , Xinran Zhang , Yu Sheng , Jianmin Ji , Yanyong Zhang

Image fusion is a technique to integrate information from multiple source images with complementary information to improve the richness of a single image. Due to insufficient task-specific training data and corresponding ground truth, most…

Computer Vision and Pattern Recognition · Computer Science 2022-01-20 Linhao Qu , Shaolei Liu , Manning Wang , Shiman Li , Siqi Yin , Qin Qiao , Zhijian Song

Motion forecasting for autonomous driving is a challenging task because complex driving scenarios result in a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Nigamaa Nayakanti , Rami Al-Rfou , Aurick Zhou , Kratarth Goel , Khaled S. Refaat , Benjamin Sapp

This paper presents a novel multi-modal Multi-Object Tracking (MOT) algorithm for self-driving cars that combines camera and LiDAR data. Camera frames are processed with a state-of-the-art 3D object detector, whereas classical clustering…

Robotics · Computer Science 2024-05-14 Riccardo Pieroni , Simone Specchia , Matteo Corno , Sergio Matteo Savaresi

Cooperative perception has been widely used in autonomous driving to alleviate the inherent limitation of single automated vehicle perception. To enable cooperation, vehicle-to-vehicle (V2V) communication plays an indispensable role. This…

Signal Processing · Electrical Eng. & Systems 2023-11-20 Chenguang Liu , Yunfei Chen , Jianjun Chen , Ryan Payton , Michael Riley , Shuang-Hua Yang

Image fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Junjia Huang , Pengxiang Yan , Jiyang Liu , Jie Wu , Zhao Wang , Yitong Wang , Liang Lin , Guanbin Li

In the typical urban intersection scenario, both vehicles and infrastructures are equipped with visual and LiDAR sensors. By successfully integrating the data from vehicle-side and road monitoring devices, a more comprehensive and accurate…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Xinyu Zhang , Yijin Xiong , Qianxin Qu , Renjie Wang , Xin Gao , Jing Liu , Shichun Guo , Jun Li

In recent times, there has been a growing focus on end-to-end autonomous driving technologies. This technology involves the replacement of the entire driving pipeline with a single neural network, which has a simpler structure and faster…

Robotics · Computer Science 2023-10-27 Hongkuan Zhou , Aifen Sui , Letian Shi , Yinxian Li

Large driving datasets are a key component in the current development and safeguarding of automated driving functions. Various methods can be used to collect such driving data records. In addition to the use of sensor equipped research…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Laurent Kloeker , Christian Geller , Amarin Kloeker , Lutz Eckstein

Multimodal remote sensing data, including spectral and lidar or photogrammetry, is crucial for achieving satisfactory land-use / land-cover classification results in urban scenes. So far, most studies have been conducted in a 2D context.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Aldino Rizaldy , Richard Gloaguen , Fabian Ewald Fassnacht , Pedram Ghamisi

This Paper proposes a novel Transformer-based end-to-end autonomous driving model named Detrive. This model solves the problem that the past end-to-end models cannot detect the position and size of traffic participants. Detrive uses an…

Robotics · Computer Science 2023-10-24 Daoming Chen , Ning Wang , Feng Chen , Tony Pipe

Fusing LiDAR and camera information is essential for achieving accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Yang Jiao , Zequn Jie , Shaoxiang Chen , Jingjing Chen , Lin Ma , Yu-Gang Jiang

Sensor fusion is crucial for an accurate and robust perception system on autonomous vehicles. Most existing datasets and perception solutions focus on fusing cameras and LiDAR. However, the collaboration between camera and radar is…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Yizhou Wang , Jen-Hao Cheng , Jui-Te Huang , Sheng-Yao Kuan , Qiqian Fu , Chiming Ni , Shengyu Hao , Gaoang Wang , Guanbin Xing , Hui Liu , Jenq-Neng Hwang

Lidars and cameras play essential roles in autonomous driving, offering complementary information for 3D detection. The state-of-the-art fusion methods integrate them at the feature level, but they mostly rely on the learned soft…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Zixuan Yin , Han Sun , Ningzhong Liu , Huiyu Zhou , Jiaquan Shen

Service mobile robots are often required to avoid dynamic objects while performing their tasks, but they usually have only limited computational resources. To further advance the practical application of service robots in complex dynamic…

Robotics · Computer Science 2026-02-25 Yushen He , Lei Zhao , Tianchen Deng , Zipeng Fang , Weidong Chen

Autonomous driving sensors generate an enormous amount of data. In this paper, we explore learned multimodal compression for autonomous driving, specifically targeted at 3D object detection. We focus on camera and LiDAR modalities and…

Image and Video Processing · Electrical Eng. & Systems 2024-08-16 Hadi Hadizadeh , Ivan V. Bajić

Estimating the 2D human poses in each view is typically the first step in calibrated multi-view 3D pose estimation. But the performance of 2D pose detectors suffers from challenging situations such as occlusions and oblique viewing angles.…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Haoyu Ma , Liangjian Chen , Deying Kong , Zhe Wang , Xingwei Liu , Hao Tang , Xiangyi Yan , Yusheng Xie , Shih-Yao Lin , Xiaohui Xie

Traffic intersections are important scenes that can be seen almost everywhere in the traffic system. Currently, most simulation methods perform well at highways and urban traffic networks. In intersection scenarios, the challenge lies in…

Robotics · Computer Science 2023-04-06 Pei Lv , Xinming Pei , Xinyu Ren , Yuzhen Zhang , Chaochao Li , Mingliang Xu

Multi-view radar-camera fused 3D object detection provides a farther detection range and more helpful features for autonomous driving, especially under adverse weather. The current radar-camera fusion methods deliver kinds of designs to…

Computer Vision and Pattern Recognition · Computer Science 2023-02-22 Zizhang Wu , Guilian Chen , Yuanzhu Gan , Lei Wang , Jian Pu

Deep learning has been used to demonstrate end-to-end neural network learning for autonomous vehicle control from raw sensory input. While LiDAR sensors provide reliably accurate information, existing end-to-end driving solutions are mainly…

Robotics · Computer Science 2021-05-21 Zhijian Liu , Alexander Amini , Sibo Zhu , Sertac Karaman , Song Han , Daniela Rus