English

Question Generation for Evaluating Cross-Dataset Shifts in Multi-modal Grounding

Computer Vision and Pattern Recognition 2022-01-25 v1

Abstract

Visual question answering (VQA) is the multi-modal task of answering natural language questions about an input image. Through cross-dataset adaptation methods, it is possible to transfer knowledge from a source dataset with larger train samples to a target dataset where training set is limited. Suppose a VQA model trained on one dataset train set fails in adapting to another, it is hard to identify the underlying cause of domain mismatch as there could exists a multitude of reasons such as image distribution mismatch and question distribution mismatch. At UCLA, we are working on a VQG module that facilitate in automatically generating OOD shifts that aid in systematically evaluating cross-dataset adaptation capabilities of VQA models.

Keywords

Cite

@article{arxiv.2201.09639,
  title  = {Question Generation for Evaluating Cross-Dataset Shifts in Multi-modal Grounding},
  author = {Arjun R. Akula},
  journal= {arXiv preprint arXiv:2201.09639},
  year   = {2022}
}
R2 v1 2026-06-24T09:00:06.251Z