English

Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge

Computer Vision and Pattern Recognition 2025-11-06 v1

Abstract

In this report, we present our solution to the MOT25-Spatiotemporal Action Grounding (MOT25-StAG) Challenge. The aim of this challenge is to accurately localize and track multiple objects that match specific and free-form language queries, using video data of complex real-world scenes as input. We model the underlying task as a video retrieval problem and present a two-stage, zero-shot approach, combining the advantages of the SOTA tracking model FastTracker and Multi-modal Large Language Model LLaVA-Video. On the MOT25-StAG test set, our method achieves m-HIoU and HOTA scores of 20.68 and 10.73 respectively, which won second place in the challenge.

Cite

@article{arxiv.2511.03332,
  title  = {Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge},
  author = {Yi Yang and Yiming Xu and Timo Kaiser and Hao Cheng and Bodo Rosenhahn and Michael Ying Yang},
  journal= {arXiv preprint arXiv:2511.03332},
  year   = {2025}
}
R2 v1 2026-07-01T07:22:38.100Z