English

Holistic Multi-modal Memory Network for Movie Question Answering

Computer Vision and Pattern Recognition 2018-11-13 v1

Abstract

Answering questions according to multi-modal context is a challenging problem as it requires a deep integration of different data sources. Existing approaches only employ partial interactions among data sources in one attention hop. In this paper, we present the Holistic Multi-modal Memory Network (HMMN) framework which fully considers the interactions between different input sources (multi-modal context, question) in each hop. In addition, it takes answer choices into consideration during the context retrieval stage. Therefore, the proposed framework effectively integrates multi-modal context, question, and answer information, which leads to more informative context retrieved for question answering. Our HMMN framework achieves state-of-the-art accuracy on MovieQA dataset. Extensive ablation studies show the importance of holistic reasoning and contributions of different attention strategies.

Keywords

Cite

@article{arxiv.1811.04595,
  title  = {Holistic Multi-modal Memory Network for Movie Question Answering},
  author = {Anran Wang and Anh Tuan Luu and Chuan-Sheng Foo and Hongyuan Zhu and Yi Tay and Vijay Chandrasekhar},
  journal= {arXiv preprint arXiv:1811.04595},
  year   = {2018}
}
R2 v1 2026-06-23T05:12:17.787Z