English

LocateBench: Evaluating the Locating Ability of Vision Language Models

Computer Vision and Pattern Recognition 2024-10-29 v1 Artificial Intelligence

Abstract

The ability to locate an object in an image according to natural language instructions is crucial for many real-world applications. In this work we propose LocateBench, a high-quality benchmark dedicated to evaluating this ability. We experiment with multiple prompting approaches, and measure the accuracy of several large vision language models. We find that even the accuracy of the strongest model, GPT-4o, lags behind human accuracy by more than 10%.

Keywords

Cite

@article{arxiv.2410.19808,
  title  = {LocateBench: Evaluating the Locating Ability of Vision Language Models},
  author = {Ting-Rui Chiang and Joshua Robinson and Xinyan Velocity Yu and Dani Yogatama},
  journal= {arXiv preprint arXiv:2410.19808},
  year   = {2024}
}

Comments

We release the dataset at https://usc-tamagotchi.github.io/locate-bench/