VALSE:以语言现象为核心的视觉与语言模型任务无关基准
计算与语言
2024-02-13 v2 计算机视觉与模式识别
摘要
我们提出 VALSE(Vision And Language Structured Evaluation,视觉与语言结构化评测),一种新颖的基准,用于测试通用预训练视觉与语言(V&L)模型在特定语言现象上的视觉-语言接地能力。VALSE 提供一套涵盖多种语言结构的六项测试。解决这些测试需要模型将语言现象在视觉模态中接地,从而实现比以往更细粒度的评测。我们使用支持构建有效干扰项(foil)的方法来建立 VALSE,并报告对五个广泛使用的 V&L 模型的评测结果。我们的实验表明,当前模型在应对大多数现象时存在相当大的困难。因此,我们期望 VALSE 能作为一个重要基准,从语言学视角衡量预训练 V&L 模型的未来进展,补充规范的以任务为中心的 V&L 评测。
引用
@article{arxiv.2112.07566,
title = {VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena},
author = {Letitia Parcalabescu and Michele Cafagna and Lilitta Muradjan and Anette Frank and Iacer Calixto and Albert Gatt},
journal= {arXiv preprint arXiv:2112.07566},
year = {2024}
}
备注
Paper accepted for publication at ACL 2022 Main; 28 pages, 4 figures, 11 tables