English

STT: Stateful Tracking with Transformers for Autonomous Driving

Robotics 2024-05-02 v1 Artificial Intelligence Computer Vision and Pattern Recognition Machine Learning

Abstract

Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTPS that address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset.

Keywords

Cite

@article{arxiv.2405.00236,
  title  = {STT: Stateful Tracking with Transformers for Autonomous Driving},
  author = {Longlong Jing and Ruichi Yu and Xu Chen and Zhengli Zhao and Shiwei Sheng and Colin Graber and Qi Chen and Qinru Li and Shangxuan Wu and Han Deng and Sangjin Lee and Chris Sweeney and Qiurui He and Wei-Chih Hung and Tong He and Xingyi Zhou and Farshid Moussavi and Zijian Guo and Yin Zhou and Mingxing Tan and Weilong Yang and Congcong Li},
  journal= {arXiv preprint arXiv:2405.00236},
  year   = {2024}
}

Comments

ICRA 2024

R2 v1 2026-06-28T16:12:19.856Z