English

Ran Score: a LLM-based Evaluation Score for Radiology Report Generation

Artificial Intelligence 2026-03-25 v1 Human-Computer Interaction

Abstract

Chest X-ray report generation and automated evaluation are limited by poor recognition of low-prevalence abnormalities and inadequate handling of clinically important language, including negation and ambiguity. We develop a clinician-guided framework combining human expertise and large language models for multi-label finding extraction from free-text chest X-ray reports and use it to define Ran Score, a finding-level metric for report evaluation. Using three non-overlapping MIMIC-CXR-EN cohorts from a public chest X-ray dataset and an independent ChestX-CN validation cohort, we optimize prompts, establish radiologist-derived reference labels and evaluate report generation models. The optimized framework improves the macro-averaged score from 0.753 to 0.956 on the MIMIC-CXR-EN development cohort, exceeds the CheXbert benchmark by 15.7 percentage points on directly comparable labels, and shows robust generalization on the ChestX-CN validation cohort. Here we show that clinician-guided prompt optimization improves agreement with a radiologist-derived reference standard and that Ran Score enables finding-level evaluation of report fidelity, particularly for low-prevalence abnormalities.

Keywords

Cite

@article{arxiv.2603.22935,
  title  = {Ran Score: a LLM-based Evaluation Score for Radiology Report Generation},
  author = {Ran Zhang and Yucong Lin and Zhaoli Su and Bowen Liu and Danni Ai and Tianyu Fu and Deqiang Xiao and Jingfan Fan and Yuanyuan Wang and Mingwei Gao and Yuwan Hu and Shuya Gao and Jingtao Li and Jian Yang and Hong Song and Hongliang Sun},
  journal= {arXiv preprint arXiv:2603.22935},
  year   = {2026}
}

Comments

4 pages, 5 figures

R2 v1 2026-07-01T11:35:01.189Z