English
Related papers

Related papers: Modality-missing RGBT Tracking: Invertible Prompt …

200 papers

Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often dominates backpropagation, leading to imbalanced optimization.…

Machine Learning · Computer Science 2025-11-19 Zhe Yang , Wenrui Li , Hongtao Chen , Penghong Wang , Ruiqin Xiong , Xiaopeng Fan

Advances in perception modeling have significantly improved the performance of object tracking. However, the current methods for specifying the target object in the initial frame are either by 1) using a box or mask template, or by 2)…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Jiawen Zhu , Zhi-Qi Cheng , Jun-Yan He , Chenyang Li , Bin Luo , Huchuan Lu , Yifeng Geng , Xuansong Xie

Camouflaged Object Detection (COD) aims to segment objects that blend seamlessly into complex backgrounds, with growing interest in exploiting additional visual modalities to enhance robustness through complementary information. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Hao Wang , Jiqing Zhang , Xin Yang , Baocai Yin , Lu Jiang , Zetian Mi , Huibing Wang

Human action recognition (HAR) with multi-modal inputs (RGB-D, skeleton, point cloud) can achieve high accuracy but typically relies on large labeled datasets and degrades sharply when sensors fail or are noisy. We present Robust…

Signal Processing · Electrical Eng. & Systems 2025-11-18 Hasan Akgul , Mari Eplik , Javier Rojas , Akira Yamamoto , Rajesh Kumar , Maya Singh

Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective context modeling, while existing context association methods based…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Yaozong Zheng , Qihua Liang , Bineng Zhong , Shuimu Zeng , Yuanliang Xue , Ning Li , Shuxiang Song

Pre-trained large multi-modal models (LMMs) exploit fine-tuning to adapt diverse user applications. Nevertheless, fine-tuning may face challenges due to deactivated sensors (e.g., cameras turned off for privacy or technical issues),…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shu Zhao , Xiaohan Zou , Tan Yu , Huijuan Xu

In this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Pengyu Zhang , Jie Zhao , Dong Wang , Huchuan Lu , Xiaoyun Yang

Multimodal data is known to be helpful for visual tracking by improving robustness to appearance variations. However, sensor synchronization challenges often compromise data availability, particularly in video settings where shortages can…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Yuedong Tan , Jiawei Shao , Eduard Zamfir , Ruanjun Li , Zhaochong An , Chao Ma , Danda Paudel , Luc Van Gool , Radu Timofte , Zongwei Wu

Salient object detection segments attractive objects in scenes. RGB and thermal modalities provide complementary information and scribble annotations alleviate large amounts of human labor. Based on the above facts, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Zhengyi Liu , Xiaoshen Huang , Guanghui Zhang , Xianyong Fang , Linbo Wang , Bin Tang

Exploiting the power of pre-trained models, prompt-based approaches stand out compared to other continual learning solutions in effectively preventing catastrophic forgetting, even with very few learnable parameters and without the need for…

Machine Learning · Computer Science 2025-01-07 Minh Le , An Nguyen , Huy Nguyen , Trang Nguyen , Trang Pham , Linh Van Ngo , Nhat Ho

The dream of achieving a student-teacher ratio of 1:1 is closer than ever thanks to the emergence of large language models (LLMs). One potential application of these models in the educational field would be to provide feedback to students…

Computers and Society · Computer Science 2025-05-06 Marc Ballestero-Ribó , Daniel Ortiz-Martínez

This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions. This is motivated by the fact that current…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Cong Hua , Qianqian Xu , Shilong Bao , Zhiyong Yang , Qingming Huang

The widespread presence of incomplete modalities in multimodal data poses a significant challenge to achieving accurate rumor detection. Existing multimodal rumor detection methods primarily focus on learning joint modality representations…

Computation and Language · Computer Science 2025-09-25 Jiajun Chen , Yangyang Wu , Xiaoye Miao , Mengying Zhu , Meng Xi

Measuring learning progress is essential for curiosity-driven exploration in reinforcement learning, but widely used signals such as prediction error often fail to distinguish meaningful, learnable patterns from random noise. This paper…

Machine Learning · Computer Science 2026-05-08 Samuel Blad , Martin Längkvist , Amy Loutfi

Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly preferred. Recently, Transformer models have shown great…

Computer Vision and Pattern Recognition · Computer Science 2022-04-13 Mengmeng Ma , Jian Ren , Long Zhao , Davide Testuggine , Xi Peng

Computer vision datasets containing multiple modalities such as color, depth, and thermal properties are now commonly accessible and useful for solving a wide array of challenging tasks. However, deploying multi-sensor heads is not possible…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Sébastien de Blois , Mathieu Garon , Christian Gagné , Jean-François Lalonde

Prompt tuning has emerged as a lightweight strategy for adapting foundation models to downstream tasks, particularly for resource-constrained systems. As pre-trained prompts become valuable assets, combining multiple source prompts offers a…

Computation and Language · Computer Science 2025-10-16 Enming Zhang , Liwen Cao , Yanru Wu , Zijie Zhao , Yang Li

The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers are still employing representations learned by the model…

Computer Vision and Pattern Recognition · Computer Science 2021-08-03 Jingxian Sun , Lichao Zhang , Yufei Zha , Abel Gonzalez-Garcia , Peng Zhang , Wei Huang , Yanning Zhang

RGB-Thermal object tracking attempt to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust feature representation learning or adaptive modal weighting.…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Rui Yang , Yabin Zhu , Xiao Wang , Chenglong Li , Jin Tang

Drone-based RGBT object detection plays a crucial role in many around-the-clock applications. However, real-world drone-viewed RGBT data suffers from the prominent position shift problem, i.e., the position of a tiny object differs greatly…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Yan Zhang , Wen Yang , Chang Xu , Qian Hu , Fang Xu , Gui-Song Xia