大型视觉语言模型是否区分了实际特征与表象特征?
摘要
人类对视觉错觉具有易感性,这些错觉对于探究感官和认知过程具有重要价值。受人类视觉研究的启发,研究者开始探索机器(如大型视觉语言模型,LVLMs)是否 exhibit similar susceptibilities to visual illusions. However, studies often have used non-abstract images and have not distinguished actual and apparent features, leading to ambiguous assessments of machine cognition. To address these limitations, we introduce a visual question answering (VQA) dataset, categorized into genuine and fake illusions, along with corresponding control images. Genuine illusions present discrepancies between actual and apparent features, whereas fake illusions have the same actual and apparent features even though they look illusory due to the similar geometric configuration. We evaluate the performance of LVLMs for genuine and fake illusion VQA tasks and investigate whether the models discern actual and apparent features. Our findings indicate that although LVLMs may appear to recognize illusions by correctly answering questions about both feature types, they predict the same answers for both Genuine Illusion and Fake Illusion VQA questions. This suggests that their responses might be based on prior knowledge of illusions rather than genuine visual understanding. The dataset is available at https://github.com/ynklab/FILM
引用
@article{arxiv.2506.05765,
title = {Do Large Vision-Language Models Distinguish between the Actual and Apparent Features of Illusions?},
author = {Taiga Shinozaki and Tomoki Doi and Amane Watahiki and Satoshi Nishida and Hitomi Yanaka},
journal= {arXiv preprint arXiv:2506.05765},
year = {2025}
}
备注
To appear in the Proceedings of the 47th Annual Meeting of the Cognitive Science Society (COGSCI 2025)