English

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

Computer Vision and Pattern Recognition 2026-03-17 v1

Abstract

Visual spatial intelligence is critical for medical image interpretation, yet remains largely unexplored in Multimodal Large Language Models (MLLMs) for 3D imaging. This gap persists due to a systemic lack of datasets featuring structured 3D spatial annotations beyond basic labels. In this study, we introduce an agentic pipeline that autonomously synthesizes spatial visual question-answering (VQA) data by orchestrating computational tools such as volume and distance calculators with multi-agent collaboration and expert radiologist validation. We present SpatialMed, the first comprehensive benchmark for evaluating 3D spatial intelligence in medical MLLMs, comprising nearly 10K question-answer pairs across multiple organs and tumor types. Our evaluations on 14 state-of-the-art MLLMs and extensive analyses reveal that current models lack robust spatial reasoning capabilities for medical imaging.

Keywords

Cite

@article{arxiv.2603.13800,
  title  = {Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space},
  author = {Quoc-Huy Trinh and Xi Ding and Yang Liu and Zhenyue Qin and Xingjian Li and Gorkem Durak and Halil Ertugrul Aktas and Elif Keles and Ulas Bagci and Min Xu},
  journal= {arXiv preprint arXiv:2603.13800},
  year   = {2026}
}