中文

ImageNetVC:面向 1000 个 ImageNet 类别的零样本与少样本视觉常识评测

计算与语言 2023-10-24 v2 人工智能 计算机视觉与模式识别

摘要

近来,大语言模型(LLMs)正作为通用接口使用,对全面视觉知识提出了显著需求。然而,当前 LLM 及其视觉增强对应模型(VaLMs)在掌握视觉常识知识方面表现如何仍不明确。为探究此问题,我们提出 ImageNetVC,一个专为零样本与少样本视觉常识评测设计、覆盖 1,000 个 ImageNet 类别的人工标注数据集。利用 ImageNetVC,我们对单模态 LLM 与 VaLMs 的基础视觉常识知识进行基准测试。此外,我们分析了影响大规模模型视觉常识知识的因素,为开发富含视觉常识知识的语言模型提供见解。我们的代码与数据集见 https://github.com/hemingkx/ImageNetVC。

关键词

引用

@article{arxiv.2305.15028,
  title  = {ImageNetVC: Zero- and Few-Shot Visual Commonsense Evaluation on 1000 ImageNet Categories},
  author = {Heming Xia and Qingxiu Dong and Lei Li and Jingjing Xu and Tianyu Liu and Ziwei Qin and Zhifang Sui},
  journal= {arXiv preprint arXiv:2305.15028},
  year   = {2023}
}

备注

EMNLP 2023 Findings (Long Paper), camera-ready version