English

Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Computer Vision and Pattern Recognition 2019-04-09 v1 Computation and Language Machine Learning

Abstract

This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to better model linguistic structure. We introduce a new data collection scheme based on grammatical constraints for surface realization to enable us to investigate the problem of grounding spatio-temporal identifying descriptions in videos. We then propose a two-stream modular attention network that learns and grounds spatio-temporal identifying descriptions based on appearance and motion. We show that motion modules help to ground motion-related words and also help to learn in appearance modules because modular neural networks resolve task interference between modules. Finally, we propose a future challenge and a need for a robust system arising from replacing ground truth visual annotations with automatic video object detector and temporal event localization.

Keywords

Cite

@article{arxiv.1904.03885,
  title  = {Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions},
  author = {Peratham Wiriyathammabhum and Abhinav Shrivastava and Vlad I. Morariu and Larry S. Davis},
  journal= {arXiv preprint arXiv:1904.03885},
  year   = {2019}
}
R2 v1 2026-06-23T08:32:32.054Z