English

Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

Computer Vision and Pattern Recognition 2022-05-06 v2

Abstract

We propose an effective two-stage approach to tackle the problem of language-based Human-centric Spatio-Temporal Video Grounding (HC-STVG) task. In the first stage, we propose an Augmented 2D Temporal Adjacent Network (Augmented 2D-TAN) to temporally ground the target moment corresponding to the given description. Primarily, we improve the original 2D-TAN from two aspects: First, a temporal context-aware Bi-LSTM Aggregation Module is developed to aggregate clip-level representations, replacing the original max-pooling. Second, we propose to employ Random Concatenation Augmentation (RCA) mechanism during the training phase. In the second stage, we use pretrained MDETR model to generate per-frame bounding boxes via language query, and design a set of hand-crafted rules to select the best matching bounding box outputted by MDETR for each frame within the grounded moment.

Keywords

Cite

@article{arxiv.2106.10634,
  title  = {Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding},
  author = {Chaolei Tan and Zihang Lin and Jian-Fang Hu and Xiang Li and Wei-Shi Zheng},
  journal= {arXiv preprint arXiv:2106.10634},
  year   = {2022}
}

Comments

Best Paper Award at the 3rd Person in Context (PIC) Challenge CVPR Workshop 2021

R2 v1 2026-06-24T03:23:45.928Z