English

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

Machine Learning 2016-06-20 v1 Computer Vision and Pattern Recognition

Abstract

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Overall, our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans.

Keywords

Cite

@article{arxiv.1606.05589,
  title  = {Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?},
  author = {Abhishek Das and Harsh Agrawal and C. Lawrence Zitnick and Devi Parikh and Dhruv Batra},
  journal= {arXiv preprint arXiv:1606.05589},
  year   = {2016}
}

Comments

5 pages, 4 figures, 3 tables, presented at 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016), New York, NY. arXiv admin note: substantial text overlap with arXiv:1606.03556