English

Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs

Computer Vision and Pattern Recognition 2026-05-05 v4 Computation and Language

Abstract

Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. In contrast, current Multimodal Large Language Models (MLLMs) are trained directly on complex downstream tasks, often bypassing these foundational visual capabilities. To systematically investigate this gap, we introduce VisFactor, a benchmark that digitizes 20 vision-centric subtests from FRCT, a well-established cognitive psychology assessment spanning four domains of human visual cognition. Furthermore, we design algorithms to automatically construct and validate unlimited test cases with controllable difficulty. Using VisFactor, we evaluate 39 frontier MLLMs, including both proprietary (e.g., GPT, Gemini) and open-source (e.g., LLaMA, Qwen) models. The best model achieves a score of only 54.0%. Analysis reveals good internal consistency (Cronbach's alpha = 0.94) and construct validity (compared to existing vision benchmarks). Models consistently fail on tasks such as mental rotation, spatial relation inference, and figure-ground discrimination, regardless of model size or prompting strategy. These findings suggest that performance improvements on existing general benchmarks might represent castles in the air instead of a genuine mastery of human-like visual cognition.

Keywords

Cite

@article{arxiv.2502.16435,
  title  = {Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs},
  author = {Jen-Tse Huang and Dasen Dai and Jen-Yuan Huang and Youliang Yuan and Xiaoyuan Liu and Wenxuan Wang and Wenxiang Jiao and Pinjia He and Zhaopeng Tu and Haodong Duan},
  journal= {arXiv preprint arXiv:2502.16435},
  year   = {2026}
}

Comments

Update: Evaluated 39 SOTA MLLMs

R2 v1 2026-06-28T21:54:21.153Z