中文

视觉与语言研究中的当前数据集综述

计算与语言 2021-08-23 v2 人工智能 计算机视觉与模式识别

摘要

融合视觉与语言长期以来一直是人工智能(AI)工作中的梦想。在过去两年中,我们见证了将视觉与语言从图像到视频及更广范围结合起来的工作爆发。可用的语料库在推进该研究领域方面发挥了关键作用。本文中,我们提出一组用于评估和分析视觉与语言数据集的质量度量,并相应地对它们分类。我们的分析表明,最新数据集已使用更复杂的语言和更抽象的概念,然而每个数据集各有不同的优势与弱点。

关键词

引用

@article{arxiv.1506.06833,
  title  = {A Survey of Current Datasets for Vision and Language Research},
  author = {Francis Ferraro and Nasrin Mostafazadeh and Ting-Hao and Huang and Lucy Vanderwende and Jacob Devlin and Michel Galley and Margaret Mitchell},
  journal= {arXiv preprint arXiv:1506.06833},
  year   = {2021}
}

备注

To appear in EMNLP 2015, short proceedings. Dataset analysis and discussion expanded, including an initial examination into reporting bias for one of them. F.F. and N.M. contributed equally to this work