We demonstrate that large multimodal language models differ substantially from humans in the distribution of coreferential expressions in a visual storytelling task. We introduce a number of metrics to quantify the characteristics of coreferential patterns in both human- and machine-written texts. Humans distribute coreferential expressions in a way that maintains consistency across texts and images, interleaving references to different entities in a highly varied way. Machines are less able to track mixed references, despite achieving perceived improvements in generation quality. Materials, metrics, and code for our study are available at https://github.com/GU-CLASP/coreference-context-scope.
@article{arxiv.2503.05298,
title = {Coreference as an indicator of context scope in multimodal narrative},
author = {Nikolai Ilinykh and Shalom Lappin and Asad Sayeed and Sharid Loáiciga},
journal= {arXiv preprint arXiv:2503.05298},
year = {2025}
}
Comments
19 pages, 4 tables. Accepted to GEM2 Workshop: Generation, Evaluation & Metrics at ACL 2025