English

AliFuse: Aligning and Fusing Multi-modal Medical Data for Computer-Aided Diagnosis

Computer Vision and Pattern Recognition 2025-02-03 v3

Abstract

Medical data collected for diagnostic decisions are typically multimodal, providing comprehensive information on a subject. While computer-aided diagnosis systems can benefit from multimodal inputs, effectively fusing such data remains a challenging task and a key focus in medical research. In this paper, we propose a transformer-based framework, called Alifuse, for aligning and fusing multimodal medical data. Specifically, we convert medical images and both unstructured and structured clinical records into vision and language tokens, employing intramodal and intermodal attention mechanisms to learn unified representations of all imaging and non-imaging data for classification. Additionally, we integrate restoration modeling with contrastive learning frameworks, jointly learning the high-level semantic alignment between images and texts and the low-level understanding of one modality with the help of another. We apply Alifuse to classify Alzheimer's disease, achieving state-of-the-art performance on five public datasets and outperforming eight baselines.

Keywords

Cite

@article{arxiv.2401.01074,
  title  = {AliFuse: Aligning and Fusing Multi-modal Medical Data for Computer-Aided Diagnosis},
  author = {Qiuhui Chen and Yi Hong},
  journal= {arXiv preprint arXiv:2401.01074},
  year   = {2025}
}

Comments

BIBM 2024

R2 v1 2026-06-28T14:06:39.364Z