English

How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People

Human-Computer Interaction 2026-02-20 v2 Artificial Intelligence

Abstract

Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer conversational assistance, where users can ask questions to obtain goal-relevant details. However, evidence about their performance in the real-world and implications for BLV people's daily lives remains limited. To address this, we conducted a two-week diary study, where we captured 20 BLV participants' use of an MLLM-enabled visual interpretation application. Although participants rated the visual interpretations of the application as "trustworthy" (mean=3.76 out of 5, max=extremely trustworthy) and "somewhat satisfying" (mean=4.13 out of 5, max=very satisfying), the AI often produced incorrect answers (22.2%) or abstained (10.8%) from responding to users' requests. Our findings show that while MLLMs can improve visual interpretations' descriptive accuracy, supporting everyday use also depends on the "visual assistant" skill: behaviors for providing goal-directed, reliable assistance. We conclude by proposing the "visual assistant" skill and guidelines to help MLLM-enabled visual interpretation applications better support BLV people's access to visual information.

Keywords

Cite

@article{arxiv.2602.13469,
  title  = {How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People},
  author = {Ricardo E. Gonzalez Penuela and Crescentia Jung and Sharon Y Lin and Ruiying Hu and Shiri Azenkot},
  journal= {arXiv preprint arXiv:2602.13469},
  year   = {2026}
}

Comments

24 pages, 17 figures, 7 tables, appendix section, to appear main track CHI 2026

R2 v1 2026-07-01T10:36:17.117Z