English

Measuring How (Not Just Whether) VLMs Build Common Ground

Computation and Language 2025-09-05 v1 Artificial Intelligence

Abstract

Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.

Keywords

Cite

@article{arxiv.2509.03805,
  title  = {Measuring How (Not Just Whether) VLMs Build Common Ground},
  author = {Saki Imai and Mert İnan and Anthony Sicilia and Malihe Alikhani},
  journal= {arXiv preprint arXiv:2509.03805},
  year   = {2025}
}
R2 v1 2026-07-01T05:20:13.648Z