English

Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

Computer Vision and Pattern Recognition 2026-08-02 v1

Abstract

Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy computational burden of state-of-the-art models hinders their deployment on resource-constrained edge devices. In this work, we identify the prediction head as a critical but often overlooked efficiency bottleneck. By strategically streamlining the decoder architecture, we unlock the potential for real-time inference but simultaneously introduce a capacity gap between the lightweight student and the heavy teacher. To resolve this, we conduct a systematic analysis of 17 distillation strategies and introduce a Dual-Alignment Distillation framework. Our key insight is that effective compression requires decoupling knowledge transfer into two complementary streams: (1) Spatial Representation Alignment, which employs feature distillation to sharpen the student's spatial focus on foreground targets ("Where to track"); and (2) Semantic Distribution Alignment, which utilizes logit-based distillation to align decision boundaries and transfer discriminative dark knowledge ("What to track"). Extensive experiments across five benchmarks demonstrate that our approach significantly outperforms complex state-of-the-art methods. Notably, our distilled model achieves 91.5% MPR on RGBT234 and operates at 54 FPS on a single RTX 4090, representing a 5x speedup over the teacher model while maintaining superior accuracy.

Cite

@article{arxiv.2608.01488,
  title  = {Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning},
  author = {Yuqi Li and Yuedong Tan and Huiran Duan and Weilun Feng and Chuanguang Yang and Zhulin An and Zongwei Wu and Shiping Wen and Tingwen Huang and Yingli Tian},
  journal= {arXiv preprint arXiv:2608.01488},
  year   = {2026}
}