English

MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise

Computer Vision and Pattern Recognition 2025-08-27 v1

Abstract

Current fundus image analysis models are predominantly built for specific tasks relying on individual datasets. The learning process is usually based on data-driven paradigm without prior knowledge, resulting in poor transferability and generalizability. To address this issue, we propose MM-Retinal, a multi-modal dataset that encompasses high-quality image-text pairs collected from professional fundus diagram books. Moreover, enabled by MM-Retinal, we present a novel Knowledge-enhanced foundational pretraining model which incorporates Fundus Image-Text expertise, called KeepFIT. It is designed with image similarity-guided text revision and mixed training strategy to infuse expert knowledge. Our proposed fundus foundation model achieves state-of-the-art performance across six unseen downstream tasks and holds excellent generalization ability in zero-shot and few-shot scenarios. MM-Retinal and KeepFIT are available at https://github.com/lxirich/MM-Retinal.

Keywords

Cite

@article{arxiv.2405.11793,
  title  = {MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise},
  author = {Ruiqi Wu and Chenran Zhang and Jianle Zhang and Yi Zhou and Tao Zhou and Huazhu Fu},
  journal= {arXiv preprint arXiv:2405.11793},
  year   = {2025}
}

Comments

Early Accepted by The International Conference on Medical Image Computing and Computer Assisted Intervention(MICCAI)2024

R2 v1 2026-06-28T16:32:44.366Z