English

A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data

Computer Vision and Pattern Recognition 2025-03-04 v1 Machine Learning

Abstract

Chest diseases rank among the most prevalent and dangerous global health issues. Object detection and phrase grounding deep learning models interpret complex radiology data to assist healthcare professionals in diagnosis. Object detection locates abnormalities for classes, while phrase grounding locates abnormalities for textual descriptions. This paper investigates how text enhances abnormality localization in chest X-rays by comparing the performance and explainability of these two tasks. To establish an explainability baseline, we proposed an automatic pipeline to generate image regions for report sentences using radiologists' eye-tracking data. The better performance - mIoU = 0.36 vs. 0.20 - and explainability - Containment ratio 0.48 vs. 0.26 - of the phrase grounding model infers the effectiveness of text in enhancing chest X-ray abnormality localization.

Keywords

Cite

@article{arxiv.2503.01037,
  title  = {A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data},
  author = {Elham Ghelichkhan and Tolga Tasdizen},
  journal= {arXiv preprint arXiv:2503.01037},
  year   = {2025}
}

Comments

Accepted in 2025 IEEE International Symposium on Biomedical Imaging (ISBI 2025)

R2 v1 2026-06-28T22:03:52.489Z