视觉问答可能仅需图像描述即可
计算机视觉与模式识别
2022-05-05 v1 计算与语言
摘要
视觉问答(VQA)已从日益复杂的模型中获益,但在数据创建方面并未享有同等的关注。本文提出一种方法,通过利用丰富的现有图像-描述标注结合用于文本问题生成的神经模型,自动大规模生成 VQA 样例。我们表明所生成的数据具有高质量。在我们的数据上训练的 VQA 模型将最先进的零样本准确率提升了两位数,并实现了在同为人工标注 VQA 数据上训练所缺乏的鲁棒性。
引用
@article{arxiv.2205.01883,
title = {All You May Need for VQA are Image Captions},
author = {Soravit Changpinyo and Doron Kukliansky and Idan Szpektor and Xi Chen and Nan Ding and Radu Soricut},
journal= {arXiv preprint arXiv:2205.01883},
year = {2022}
}
备注
2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2022)