English

Recent Advances in Video Question Answering: A Review of Datasets and Methods

Computer Vision and Pattern Recognition 2021-03-19 v2

Abstract

Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation have preceded the task of VQA. VQA helps to retrieve temporal and spatial information from the video scenes and interpret it. In this survey, we review a number of methods and datasets for the task of VQA. To the best of our knowledge, no previous survey has been conducted for the VQA task.

Keywords

Cite

@article{arxiv.2101.05954,
  title  = {Recent Advances in Video Question Answering: A Review of Datasets and Methods},
  author = {Devshree Patel and Ratnam Parikh and Yesha Shastri},
  journal= {arXiv preprint arXiv:2101.05954},
  year   = {2021}
}

Comments

18 pages, 5 tables, Video and Image Question Answering Workshop, 25th International Conference on Pattern Recognition