English

An Analysis of Action Recognition Datasets for Language and Vision Tasks

Computation and Language 2017-04-25 v1 Computer Vision and Pattern Recognition

Abstract

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene under- standing and image retrieval. In this survey, we categorize the existing ap- proaches based on how they conceptualize this problem and provide a detailed review of existing datasets, highlighting their di- versity as well as advantages and disad- vantages. We focus on recently devel- oped datasets which link visual informa- tion with linguistic resources and provide a fine-grained syntactic and semantic anal- ysis of actions in images.

Keywords

Cite

@article{arxiv.1704.07129,
  title  = {An Analysis of Action Recognition Datasets for Language and Vision Tasks},
  author = {Spandana Gella and Frank Keller},
  journal= {arXiv preprint arXiv:1704.07129},
  year   = {2017}
}

Comments

To appear in Proceedings of ACL 2017, 8 pages

R2 v1 2026-06-22T19:25:28.869Z