English

Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking

Computer Vision and Pattern Recognition 2020-07-07 v1

Abstract

In this study, we propose a novel RGB-T tracking framework by jointly modeling both appearance and motion cues. First, to obtain a robust appearance model, we develop a novel late fusion method to infer the fusion weight maps of both RGB and thermal (T) modalities. The fusion weights are determined by using offline-trained global and local multimodal fusion networks, and then adopted to linearly combine the response maps of RGB and T modalities. Second, when the appearance cue is unreliable, we comprehensively take motion cues, i.e., target and camera motions, into account to make the tracker robust. We further propose a tracker switcher to switch the appearance and motion trackers flexibly. Numerous results on three recent RGB-T tracking datasets show that the proposed tracker performs significantly better than other state-of-the-art algorithms.

Keywords

Cite

@article{arxiv.2007.02041,
  title  = {Jointly Modeling Motion and Appearance Cues for Robust RGB-T Tracking},
  author = {Pengyu Zhang and Jie Zhao and Dong Wang and Huchuan Lu and Xiaoyun Yang},
  journal= {arXiv preprint arXiv:2007.02041},
  year   = {2020}
}

Comments

12 pages, 11 figures

R2 v1 2026-06-23T16:50:55.525Z