English

Video Question Generation via Cross-Modal Self-Attention Networks Learning

Computer Vision and Pattern Recognition 2020-02-18 v3 Computation and Language

Abstract

We introduce a novel task, Video Question Generation (Video QG). A Video QG model automatically generates questions given a video clip and its corresponding dialogues. Video QG requires a range of skills -- sentence comprehension, temporal relation, the interplay between vision and language, and the ability to ask meaningful questions. To address this, we propose a novel semantic rich cross-modal self-attention (SRCMSA) network to aggregate the multi-modal and diverse features. To be more precise, we enhance the video frames semantic by integrating the object-level information, and we jointly consider the cross-modal attention for the video question generation task. Excitingly, our proposed model remarkably improves the baseline from 7.58 to 14.48 in the BLEU-4 score on the TVQA dataset. Most of all, we arguably pave a novel path toward understanding the challenging video input and we provide detailed analysis in terms of diversity, which ushers the avenues for future investigations.

Keywords

Cite

@article{arxiv.1907.03049,
  title  = {Video Question Generation via Cross-Modal Self-Attention Networks Learning},
  author = {Yu-Siang Wang and Hung-Ting Su and Chen-Hsi Chang and Zhe-Yu Liu and Winston H. Hsu},
  journal= {arXiv preprint arXiv:1907.03049},
  year   = {2020}
}

Comments

Accepted by ICASSP 2020

R2 v1 2026-06-23T10:13:40.297Z