ImageNetVC:面向 1000 个 ImageNet 类别的零样本与少样本视觉常识评测
计算与语言
2023-10-24 v2 人工智能
计算机视觉与模式识别
摘要
近来,大语言模型(LLMs)正作为通用接口使用,对全面视觉知识提出了显著需求。然而,当前 LLM 及其视觉增强对应模型(VaLMs)在掌握视觉常识知识方面表现如何仍不明确。为探究此问题,我们提出 ImageNetVC,一个专为零样本与少样本视觉常识评测设计、覆盖 1,000 个 ImageNet 类别的人工标注数据集。利用 ImageNetVC,我们对单模态 LLM 与 VaLMs 的基础视觉常识知识进行基准测试。此外,我们分析了影响大规模模型视觉常识知识的因素,为开发富含视觉常识知识的语言模型提供见解。我们的代码与数据集见 https://github.com/hemingkx/ImageNetVC。
引用
@article{arxiv.2305.15028,
title = {ImageNetVC: Zero- and Few-Shot Visual Commonsense Evaluation on 1000 ImageNet Categories},
author = {Heming Xia and Qingxiu Dong and Lei Li and Jingjing Xu and Tianyu Liu and Ziwei Qin and Zhifang Sui},
journal= {arXiv preprint arXiv:2305.15028},
year = {2023}
}
备注
EMNLP 2023 Findings (Long Paper), camera-ready version