LocateBench: 评估视觉语言模型定位能力的基准
计算机视觉与模式识别
2024-10-29 v1 人工智能
摘要
根据自然语言指令在图像中定位对象是许多实际应用中至关重要的能力。在本工作中,我们提出了 LocateBench, 一个高质量的基准,专用于评估这一能力。我们实验了多种提示方法,并衡量了几个大型视觉语言模型的准确率。我们发现,即使是最强的模型 GPT-4o 的准确率也比人类准确率差得超过 10%。
引用
@article{arxiv.2410.19808,
title = {LocateBench: Evaluating the Locating Ability of Vision Language Models},
author = {Ting-Rui Chiang and Joshua Robinson and Xinyan Velocity Yu and Dani Yogatama},
journal= {arXiv preprint arXiv:2410.19808},
year = {2024}
}
备注
We release the dataset at https://usc-tamagotchi.github.io/locate-bench/