中文
相关论文

相关论文: RGBT Tracking via Multi-Adapter Network with Hiera…

200 篇论文

In this paper, we explore adapter tuning and introduce a novel dual-adapter architecture for spatio-temporal multimodal tracking, dubbed DMTrack. The key of our DMTrack lies in two simple yet effective modules, including a spatio-temporal…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Weihong Li , Shaohua Dong , Haonan Lu , Yanhao Zhang , Heng Fan , Libo Zhang

Graph-based models have emerged as a powerful paradigm for modeling multimodal urban data and learning region representations for various downstream tasks. However, existing approaches face two major limitations. (1) They typically employ…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yaya Zhao , Kaiqi Zhao , Zixuan Tang , Zhiyuan Liu , Xiaoling Lu , Yalei Du

Existing cross-modal pedestrian detection (CMPD) employs complementary information from RGB and thermal-infrared (TIR) modalities to detect pedestrians in 24h-surveillance systems.RGB captures rich pedestrian details under daylight, while…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Qian Bie , Xiao Wang , Bin Yang , Zhixi Yu , Jun Chen , Xin Xu

Human behavior anomaly detection aims to identify unusual human actions, playing a crucial role in intelligent surveillance and other areas. The current mainstream methods still adopt reconstruction or future frame prediction techniques.…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Guoqing Yang , Zhiming Luo , Jianzhe Gao , Yingxin Lai , Kun Yang , Yifan He , Shaozi Li

RGB-thermal salient object detection (RGB-T SOD) aims to locate the common prominent objects of an aligned visible and thermal infrared image pair and accurately segment all the pixels belonging to those objects. It is promising in…

计算机视觉与模式识别 · 计算机科学 2022-07-11 Xiurong Jiang , Lin Zhu , Yifan Hou , Hui Tian

Depth completion aims to recover a dense depth map from the sparse depth data and the corresponding single RGB image. The observed pixels provide the significant guidance for the recovery of the unobserved pixels' depth. However, due to the…

计算机视觉与模式识别 · 计算机科学 2021-06-09 Shanshan Zhao , Mingming Gong , Huan Fu , Dacheng Tao

Multi-modality of color and depth, i.e., RGB-D, is of great importance in recent research of indoor scene recognition. In this kind of data representation, depth map is able to describe the 3D structure of scenes and geometric relations…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Qiong Liu , Ruofei Xiong , Xingzhen Chen , Muyao Peng , You Yang

Most modern multi-object tracking (MOT) systems follow the tracking-by-detection paradigm. It first localizes the objects of interest, then extracting their individual appearance features to make data association. The individual features,…

计算机视觉与模式识别 · 计算机科学 2020-07-02 Tianyi Liang , Long Lan , Zhigang Luo

Depth estimation in complex real-world scenarios is a challenging task, especially when relying solely on a single modality such as visible light or thermal infrared (THR) imagery. This paper proposes a novel multimodal depth estimation…

图像与视频处理 · 电气工程与系统科学 2025-04-30 Zelin Meng , Takanori Fukao

The process of association and tracking of sensor detections is a key element in providing situational awareness. When the targets in the scenario are dense and exhibit high maneuverability, Multi-Target Tracking (MTT) becomes a challenging…

机器学习 · 计算机科学 2020-11-20 Rishabh Verma , R Rajesh , MS Easwaran

Recent works in multiple object tracking use sequence model to calculate the similarity score between the detections and the previous tracklets. However, the forced exposure to ground-truth in the training stage leads to the…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Tao Hu , Lichao Huang , Han Shen

Multi-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt fixed fusion structures to integrate multi-modal feature,…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Andong Lu , Wanyu Wang , Chenglong Li , Jin Tang , Bin Luo

Integrating different representations from complementary sensing modalities is crucial for robust scene interpretation in autonomous driving. While deep learning architectures that fuse vision and range data for 2D object detection have…

计算机视觉与模式识别 · 计算机科学 2022-03-08 George Eskandar , Robert A. Marsden , Pavithran Pandiyan , Mario Döbler , Karim Guirguis , Bin Yang

Existing RGB-Thermal Video Object Detection (RGBT VOD) methods predominantly rely on the manual alignment of image pairs, that is both labor-intensive and time-consuming. This dependency significantly restricts the scalability and practical…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Qishun Wang , Zhengzheng Tu , Kunpeng Wang , Le Gu , Chuanwang Guo

Recent progresses in visual tracking have greatly improved the tracking performance. However, challenges such as occlusion and view change remain obstacles in real world deployment. A natural solution to these challenges is to use multiple…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Minye Wu , Haibin Ling , Ning Bi , Shenghua Gao , Hao Sheng , Jingyi Yu

The problem of multi-object tracking is a fundamental computer vision research focus, widely used in public safety, transport, autonomous vehicles, robotics, and other regions involving artificial intelligence. Because of the complexity of…

计算机视觉与模式识别 · 计算机科学 2022-10-20 Kai Ren , Chuanping Hu

In the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a black box. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Hao Chen , Haoran Zhou , Yunshu Zhang , Zheng Lin , Yongjian Deng

Integrating multispectral data in object detection, especially visible and infrared images, has received great attention in recent years. Since visible (RGB) and infrared (IR) images can provide complementary information to handle light…

计算机视觉与模式识别 · 计算机科学 2022-09-29 Maoxun Yuan , Yinyan Wang , Xingxing Wei

Action recognition has been a heated topic in computer vision for its wide application in vision systems. Previous approaches achieve improvement by fusing the modalities of the skeleton sequence and RGB video. However, such methods have a…

计算机视觉与模式识别 · 计算机科学 2022-02-24 Xiaoguang Zhu , Ye Zhu , Haoyu Wang , Honglin Wen , Yan Yan , Peilin Liu

Gesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to…

计算机视觉与模式识别 · 计算机科学 2021-07-07 Zitong Yu , Benjia Zhou , Jun Wan , Pichao Wang , Haoyu Chen , Xin Liu , Stan Z. Li , Guoying Zhao