中文

视觉问答中的人类注意力:人类与深度网络关注相同区域吗?

机器学习 2016-06-20 v1 计算机视觉与模式识别

摘要

我们对视觉问答 (VQA) 中的“人类注意力”进行了大规模研究,以理解人类在回答关于图像的问题时选择看向何处。我们设计并测试了多种受游戏启发的新型注意力标注界面,这些界面要求受试者锐化模糊图像中的区域以回答问题。由此,我们引入了 VQA-HAT(人类注意力)数据集。我们通过定性(可视化)和定量(秩序相关)两种方式,将最先进的 VQA 模型生成的注意力图与人类注意力进行了比较评估。总体而言,我们的实验表明,当前 VQA 中的注意力模型似乎并未关注与人类相同的区域。

关键词

引用

@article{arxiv.1606.05589,
  title  = {Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?},
  author = {Abhishek Das and Harsh Agrawal and C. Lawrence Zitnick and Devi Parikh and Dhruv Batra},
  journal= {arXiv preprint arXiv:1606.05589},
  year   = {2016}
}

备注

5 pages, 4 figures, 3 tables, presented at 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016), New York, NY. arXiv admin note: substantial text overlap with arXiv:1606.03556