English

Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features

Computer Vision and Pattern Recognition 2023-01-12 v2

Abstract

Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detection utilizing pretrained image-text co-embeddings. Despite being trained on static images rather than videos, we show that image-text co-embeddings enable openvocabulary performance competitive with fully-supervised models. We show that the performance can be further improved by ensembling the image-text features with features encoding local motion, like optical flow based features, or other modalities, like audio. In addition, we propose a more reasonable open-vocabulary evaluation setting for the ActivityNet data set, where the category splits are based on similarity rather than random assignment.

Keywords

Cite

@article{arxiv.2212.10596,
  title  = {Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features},
  author = {Vivek Rathod and Bryan Seybold and Sudheendra Vijayanarasimhan and Austin Myers and Xiuye Gu and Vighnesh Birodkar and David A. Ross},
  journal= {arXiv preprint arXiv:2212.10596},
  year   = {2023}
}
R2 v1 2026-06-28T07:45:34.722Z