English

Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports

Artificial Intelligence 2026-04-22 v1

Abstract

Accurate disease classification from radiology reports is essential for many applications. While supervised fine-tuning (SFT) of lightweight LLMs improves accuracy, it can degrade reasoning. We propose a two-stage approach: SFT on disease labels followed by Group Relative Policy Optimization (GRPO) to refine predictions by optimizing accuracy and format without reasoning supervision. Across three radiologist-annotated datasets, SFT outperformed baselines and GRPO further improved classification and enhanced reasoning recall and comprehensiveness.

Keywords

Cite

@article{arxiv.2604.19060,
  title  = {Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports},
  author = {Yishu Wei and Yi Lin and Adam Flanders and George Shih and Yifan Peng},
  journal= {arXiv preprint arXiv:2604.19060},
  year   = {2026}
}