English

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

Computer Vision and Pattern Recognition 2026-07-06 v1

Abstract

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources. In this paper, we study this missing step, i.e., raw medical data standardization. Specifically, models are given raw dataset folders and evaluated on their ability to identify source formats, convert raw medical images into VLM-compatible visual inputs, extract relevant textual information, and organize the results into structured image-text pairs. To construct this Medical Data Standardization Benchmark (MDS-Bench), we manually annotate 1,939 raw medical data standardization tasks covering diverse clinical practice, radiology modalities, annotation formats, and directory layouts. Extensive experiments show that even the best performing VLMs, i.e., Gemini 3 Flash, achieve only 48.6% end-to-end success rate. Our research highlights raw medical data standardization as a critical bottleneck for medical AI diagnosis in real practice.

Cite

@article{arxiv.2607.04694,
  title  = {Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?},
  author = {Xin Chen and Dongliang Xu and Cunhao Zhu and Xudong Luo and Haoyang Lyu and Xiaoxiao Sun and Serena Yeung-Levy and Yue Yao},
  journal= {arXiv preprint arXiv:2607.04694},
  year   = {2026}
}

Comments

16 pages, 7 figures