English

Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images

Computer Vision and Pattern Recognition 2023-06-16 v3

Abstract

While multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nature of medical tasks and the high demand for domain knowledge. To address this challenge, we propose a novel approach called Knowledge-enhanced Auto Diagnosis (KAD) which leverages existing medical domain knowledge to guide vision-language pre-training using paired chest X-rays and radiology reports. We evaluate KAD on {four} external X-ray datasets and demonstrate that its zero-shot performance is not only comparable to that of fully-supervised models, but also superior to the average of three expert radiologists for three (out of five) pathologies with statistical significance. Moreover, when few-shot annotation is available, KAD outperforms all existing approaches in fine-tuning settings, demonstrating its potential for application in different clinical scenarios.

Keywords

Cite

@article{arxiv.2302.14042,
  title  = {Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images},
  author = {Xiaoman Zhang and Chaoyi Wu and Ya Zhang and Yanfeng Wang and Weidi Xie},
  journal= {arXiv preprint arXiv:2302.14042},
  year   = {2023}
}
R2 v1 2026-06-28T08:50:57.287Z