English

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

Artificial Intelligence 2026-04-10 v2 Computer Vision and Pattern Recognition

Abstract

Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathematics, physics and chemistry, organized into three cognitive levels: perception and recognition, combination and reasoning, and association and critical thinking. Across leading MLLMs, we observe a consistent cognitive mismatch. Models frequently underperform on elementary symbol recognition while appearing relatively competent on more complex reasoning tasks. This recognition-reasoning inversion indicates that current systems often compensate with linguistic priors, template retrieval or procedural reasoning instead of robust visual grounding. The pattern is especially clear for sparse, low-redundancy symbols such as handwritten characters, formula graphs, circuit diagrams and chemical structures. These results show that symbolic understanding remains a major bottleneck for multimodal intelligence and motivate training and evaluation schemes that prioritize grounded perception in discrete semantic spaces.

Keywords

Cite

@article{arxiv.2603.18472,
  title  = {Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding},
  author = {Yinghui Li and Jiayi Kuang and Peng Xing and Daixian Liu and Yongheng Zhang and Junnan Dong and Shu-Yu Guo and Yangning Li and Qingyu Zhou and Wenhao Jiang and Hai-Tao Zheng and Ying Shen and Liang Lin and Philip S. Yu},
  journal= {arXiv preprint arXiv:2603.18472},
  year   = {2026}
}
R2 v1 2026-07-01T11:27:26.833Z