English

All You May Need for VQA are Image Captions

Computer Vision and Pattern Recognition 2022-05-05 v1 Computation and Language

Abstract

Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-caption annotations combined with neural models for textual question generation. We show that the resulting data is of high-quality. VQA models trained on our data improve state-of-the-art zero-shot accuracy by double digits and achieve a level of robustness that lacks in the same model trained on human-annotated VQA data.

Keywords

Cite

@article{arxiv.2205.01883,
  title  = {All You May Need for VQA are Image Captions},
  author = {Soravit Changpinyo and Doron Kukliansky and Idan Szpektor and Xi Chen and Nan Ding and Radu Soricut},
  journal= {arXiv preprint arXiv:2205.01883},
  year   = {2022}
}

Comments

2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2022)

R2 v1 2026-06-24T11:06:41.179Z