English

One Graph to Track Them All: Dynamic GNNs for Single- and Multi-View Tracking

Computer Vision and Pattern Recognition 2026-01-01 v2

Abstract

This work presents a unified, fully differentiable model for multi-people tracking that learns to associate detections into trajectories without relying on pre-computed tracklets. The model builds a dynamic spatiotemporal graph that aggregates spatial, contextual, and temporal information, enabling seamless information propagation across entire sequences. To improve occlusion handling, the graph can also encode scene-specific information. We also introduce a new large-scale dataset with 25 partially overlapping views, detailed scene reconstructions, and extensive occlusions. Experiments show the model achieves state-of-the-art performance on public benchmarks and the new dataset, with flexibility across diverse conditions. Both the dataset and approach will be publicly released to advance research in multi-people tracking.

Keywords

Cite

@article{arxiv.2507.08494,
  title  = {One Graph to Track Them All: Dynamic GNNs for Single- and Multi-View Tracking},
  author = {Martin Engilberge and Ivan Vrkic and Friedrich Wilke Grosche and Julien Pilet and Engin Turetken and Pascal Fua},
  journal= {arXiv preprint arXiv:2507.08494},
  year   = {2026}
}
R2 v1 2026-07-01T03:56:25.138Z