English

Where To Look: Focus Regions for Visual Question Answering

Computer Vision and Pattern Recognition 2016-01-12 v2

Abstract

We present a method that learns to answer visual questions by selecting image regions relevant to the text-based query. Our method exhibits significant improvements in answering questions such as "what color," where it is necessary to evaluate a specific location, and "what room," where it selectively identifies informative image regions. Our model is tested on the VQA dataset which is the largest human-annotated visual question answering dataset to our knowledge.

Keywords

Cite

@article{arxiv.1511.07394,
  title  = {Where To Look: Focus Regions for Visual Question Answering},
  author = {Kevin J. Shih and Saurabh Singh and Derek Hoiem},
  journal= {arXiv preprint arXiv:1511.07394},
  year   = {2016}
}

Comments

Submitted to CVPR2016

R2 v1 2026-06-22T11:52:26.995Z