English

TrajTok: Learning Trajectory Tokens enables better Video Understanding

Computer Vision and Pattern Recognition 2026-05-12 v2

Abstract

Tokenization in video models, typically through patchification, generates an excessive and redundant number of tokens. This severely limits video efficiency and scalability. While recent trajectory-based tokenizers offer a promising solution by decoupling video duration from token count, they rely on complex external segmentation and tracking pipelines that are slow and task-agnostic. We propose TrajTok, an end-to-end video tokenizer module that is fully integrated and co-trained with video models for a downstream objective, dynamically adapting its token granularity to semantic complexity, independent of video duration. TrajTok contains a unified segmenter that performs implicit clustering over pixels in both space and time to directly produce object trajectories in a single forward pass. By prioritizing downstream adaptability over pixel-perfect segmentation fidelity, TrajTok is lightweight and efficient, yet empirically improves video understanding performance. With TrajTok, we implement a video CLIP model trained from scratch (TrajViT2). It achieves the best accuracy at scale across both classification and retrieval benchmarks, while maintaining efficiency comparable to the best token-merging methods. TrajTok also proves to be a versatile component beyond its role as a tokenizer. We show that it can be seamlessly integrated as either a probing head for pretrained visual features (TrajAdapter) or an alignment connector in vision-language models (TrajVLM) with especially strong performance in long-video reasoning.

Keywords

Cite

@article{arxiv.2602.22779,
  title  = {TrajTok: Learning Trajectory Tokens enables better Video Understanding},
  author = {Chenhao Zheng and Jieyu Zhang and Jianing Zhang and Weikai Huang and Ashutosh Kumar and Quan Kong and Oncel Tuzel and Chun-Liang Li and Ranjay Krishna},
  journal= {arXiv preprint arXiv:2602.22779},
  year   = {2026}
}

Comments

CVPR 2026

R2 v1 2026-07-01T10:53:33.186Z