English

Vision-Language Models under Cultural and Inclusive Considerations

Computer Vision and Pattern Recognition 2024-07-09 v1 Artificial Intelligence Computation and Language Computers and Society

Abstract

Large vision-language models (VLMs) can assist visually impaired people by describing images from their daily lives. Current evaluation datasets may not reflect diverse cultural user backgrounds or the situational context of this use case. To address this problem, we create a survey to determine caption preferences and propose a culture-centric evaluation benchmark by filtering VizWiz, an existing dataset with images taken by people who are blind. We then evaluate several VLMs, investigating their reliability as visual assistants in a culturally diverse setting. While our results for state-of-the-art models are promising, we identify challenges such as hallucination and misalignment of automatic evaluation metrics with human judgment. We make our survey, data, code, and model outputs publicly available.

Keywords

Cite

@article{arxiv.2407.06177,
  title  = {Vision-Language Models under Cultural and Inclusive Considerations},
  author = {Antonia Karamolegkou and Phillip Rust and Yong Cao and Ruixiang Cui and Anders Søgaard and Daniel Hershcovich},
  journal= {arXiv preprint arXiv:2407.06177},
  year   = {2024}
}

Comments

HuCLLM @ ACL 2024

R2 v1 2026-06-28T17:33:15.612Z