English

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

Computer Vision and Pattern Recognition 2026-03-04 v2 Artificial Intelligence

Abstract

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision-language alignment and (2) cross-modal feature fusion. In this work, we propose MedMAP, a Medical Modality-Aware Pretraining framework that enhances vision-language representation learning in 3D MRI. MedMAP comprises a modality-aware vision-language alignment stage and a fine-tuning stage for multi-organ abnormality detection. During the pre-training stage, the modality-aware encoders implicitly capture the joint modality distribution and improve alignment between visual and textual representations. We then fine-tune the pre-trained vision encoders (while keeping the text encoder frozen) for downstream tasks. To this end, we curated MedMoM-MRI3D, comprising 7,392 3D MRI volume-report pairs spanning twelve MRI modalities and nine abnormalities tailored for various 3D medical analysis tasks. Extensive experiments on MedMoM-MRI3D demonstrate that MedMAP significantly outperforms existing VLMs in 3D MRI-based multi-organ abnormality detection. Our code is available at https://github.com/RomantiDr/MedMAP.

Keywords

Cite

@article{arxiv.2602.23652,
  title  = {3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection},
  author = {Haowen Zhu and Ning Yin and Xiaogen Zhou},
  journal= {arXiv preprint arXiv:2602.23652},
  year   = {2026}
}
R2 v1 2026-07-01T10:54:52.542Z