English
Related papers

Related papers: Multi-Modal Fusion for End-to-End RGB-T Tracking

200 papers

Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through leveraging complementary cues from different sources, it remains…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Andong Lu , Mai Wen , Jinhu Wang , Yuanzhi Guo , Chenglong Li , Jin Tang , Bin Luo

Detecting hidden or partially concealed objects remains a fundamental challenge in multimodal environments, where factors like occlusion, camouflage, and lighting variations significantly hinder performance. Traditional RGB-based detection…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Harris Song , Tuan-Anh Vu , Sanjith Menon , Sriram Narasimhan , M. Khalid Jawed

In this work, we propose a novel staged depthwise correlation and feature fusion network, named DCFFNet, to further optimize the feature extraction for visual tracking. We build our deep tracker upon a siamese network architecture, which is…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Dianbo Ma , Jianqiang Xiao , Ziyan Gao , Satoshi Yamane

Low-quality modalities contain not only a lot of noisy information but also some discriminative features in RGBT tracking. However, the potentials of low-quality modalities are not well explored in existing RGBT tracking algorithms. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-05-02 Andong Lu , Cun Qian , Chenglong Li , Jin Tang , Liang Wang

In this paper, we propose a three-stream adaptive fusion network named TAFNet, which uses paired RGB and thermal images for crowd counting. Specifically, TAFNet is divided into one main stream and two auxiliary streams. We combine a pair of…

Computer Vision and Pattern Recognition · Computer Science 2022-02-18 Haihan Tang , Yi Wang , Lap-Pui Chau

To properly assist humans in their needs, human activity recognition (HAR) systems need the ability to fuse information from multiple modalities. Our hypothesis is that multimodal sensors, visual and non-visual tend to provide complementary…

Computer Vision and Pattern Recognition · Computer Science 2022-11-09 Hyeongju Choi , Apoorva Beedu , Harish Haresamudram , Irfan Essa

3D multi-object tracking (MOT) and trajectory forecasting are two critical components in modern 3D perception systems. We hypothesize that it is beneficial to unify both tasks under one framework to learn a shared feature representation of…

Computer Vision and Pattern Recognition · Computer Science 2020-08-27 Xinshuo Weng , Ye Yuan , Kris Kitani

Existing cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems.RGB captures rich pedestrian details under daylight, while…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Qian Bie , Xiao Wang , Bin Yang , Zhixi Yu , Jun Chen , Xin Xu

Accurate 3D multi-object tracking (MOT) is crucial for autonomous driving, as it enables robust perception, navigation, and planning in complex environments. While deep learning-based solutions have demonstrated impressive 3D MOT…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Guanhua Ding , Yuxuan Xia , Runwei Guan , Qinchen Wu , Tao Huang , Weiping Ding , Jinping Sun , Guoqiang Mao

In this paper we present a new approach for feature fusion between RGB and LWIR Thermal images for the task of semantic segmentation for driving perception. We propose DooDLeNet, a double DeepLab architecture with specialized…

Machine Learning · Computer Science 2022-04-22 Oriel Frigo , Lucien Martin-Gaffé , Catherine Wacongne

The 3D scene understanding is mainly considered as a crucial requirement in computer vision and robotics applications. One of the high-level tasks in 3D scene understanding is semantic segmentation of RGB-Depth images. With the availability…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Fahimeh Fooladgar , Shohreh Kasaei

Hyperspectral imagery encodes rich material properties that can improve tracking robustness under appearance ambiguity, illumination change, and background clutter. However, due to the limited availability of hyperspectral video data, many…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Xu Han , Mohammad Aminul Islam , Lei Wang , Zekun Long , Guanmanyi Fu , Wangshu Cai , Kuldip K. Paliwal , Jun Zhou

This study aims to improve the performance and generalization capability of end-to-end autonomous driving with scene understanding leveraging deep learning and multimodal sensor fusion techniques. The designed end-to-end deep neural network…

Robotics · Computer Science 2020-08-04 Zhiyu Huang , Chen Lv , Yang Xing , Jingda Wu

Thermal infrared (TIR) images typically lack detailed features and have low contrast, making it challenging for conventional feature extraction models to capture discriminative target characteristics. As a result, trackers are often…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Ruoyan Xiong , Huanbin Zhang , Shentao Wang , Hui He , Yuke Hou , Yue Zhang , Yujie Cui , Huipan Guan , Shang Zhang

Autonomous racing has rapidly gained research attention. Traditionally, racing cars rely on 2D LiDAR as their primary visual system. In this work, we explore the integration of an event camera with the existing system to provide enhanced…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Zhuyun Zhou , Zongwei Wu , Florian Bolli , Rémi Boutteau , Fan Yang , Radu Timofte , Dominique Ginhac , Tobi Delbruck

Multispectral pedestrian detection has gained significant attention in recent years, particularly in autonomous driving applications. To address the challenges posed by adversarial illumination conditions, the combination of thermal and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Arunkumar Rathinam , Leo Pauly , Abd El Rahman Shabayek , Wassim Rharbaoui , Anis Kacem , Vincent Gaudillière , Djamila Aouada

End-to-end autonomous driving is typically built upon imitation learning (IL), yet its performance is constrained by the quality of human demonstrations. To overcome this limitation, recent methods incorporate reinforcement learning (RL)…

Robotics · Computer Science 2026-04-13 Zhexi Lian , Haoran Wang , Xuerun Yan , Weimeng Lin , Xianhong Zhang , Yongyu Chen , Jia Hu

Tracking any point (TAP) is a fundamental yet challenging task in computer vision, requiring high precision and long-term motion reasoning. Recent attempts to combine RGB frames and event streams have shown promise, yet they typically rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jiaxiong Liu , Zhen Tan , Jinpu Zhang , Yi Zhou , Hui Shen , Xieyuanli Chen , Dewen Hu

Pedestrian detection is a critical task in robot perception. Multispectral modalities (visible light and thermal) can boost pedestrian detection performance by providing complementary visual information. Several gaps remain with…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Asiegbu Miracle Kanu-Asiegbu , Nitin Jotwani , Xiaoxiao Du

Multi-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the problem of the neglect of inter-feature relationships, high-frequency…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Xiaoli Zhang , Liying Wang , Libo Zhao , Xiongfei Li , Siwei Ma