English
Related papers

Related papers: Radiology-Aware Model-Based Evaluation Metric for …

200 papers

Automated radiology report drafting (ARRD) using vision-language models (VLMs) has advanced rapidly, yet most systems lack explicit uncertainty estimates, limiting trust and safe clinical deployment. We propose CONRep, a model-agnostic…

The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers of reports. In this…

Computation and Language · Computer Science 2024-08-13 Alyssa Huang , Oishi Banerjee , Kay Wu , Eduardo Pontes Reis , Pranav Rajpurkar

We propose and demonstrate a novel machine learning algorithm that assesses pulmonary edema severity from chest radiographs. While large publicly available datasets of chest radiographs and free-text radiology reports exist, only limited…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Geeticka Chauhan , Ruizhi Liao , William Wells , Jacob Andreas , Xin Wang , Seth Berkowitz , Steven Horng , Peter Szolovits , Polina Golland

Automatic radiology report generation is challenging as medical images or reports are usually similar to each other due to the common content of anatomy. This makes a model hard to capture the uniqueness of individual images and is prone to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Bhanu Prakash Voutharoja , Lei Wang , Luping Zhou

Mammography report generation is a critical yet underexplored task in medical AI, characterized by challenges such as multiview image reasoning, high-resolution visual cues, and unstructured radiologic language. In this work, we introduce…

Image and Video Processing · Electrical Eng. & Systems 2025-08-14 Nak-Jun Sung , Donghyun Lee , Bo Hwa Choi , Chae Jung Park

AI-driven models have shown great promise in detecting errors in radiology reports, yet the field lacks a unified benchmark for rigorous evaluation of error detection and further correction. To address this gap, we introduce CorBenchX, a…

Artificial Intelligence · Computer Science 2025-05-20 Jing Zou , Qingqiu Li , Chenyu Lian , Lihao Liu , Xiaohan Yan , Shujun Wang , Jing Qin

Automated radiology report generation holds immense potential to alleviate the heavy workload of radiologists. Despite the formidable vision-language capabilities of recent Multimodal Large Language Models (MLLMs), their clinical deployment…

Artificial Intelligence · Computer Science 2026-03-17 Tuoshi Qi , Shenshen Bu , Yingfei Xiang , Zhiming Dai

Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease…

Tissues and Organs · Quantitative Biology 2026-04-13 Aishik Konwer , Moinak Bhattacharya , Prateek Prasanna

Automated generation of clinically accurate radiology reports can improve patient care. Previous report generation methods that rely on image captioning models often generate incoherent and incorrect text due to their lack of relevant…

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, generating plausible yet…

Computation and Language · Computer Science 2026-02-05 Ruixiao Yang , Yuanhe Tian , Xu Yang , Huiqi Li , Yan Song

Vehicle models have a long history of research and as of today are able to model the involved physics in a reasonable manner. However, each new vehicle has its new characteristics or parameters. The identification of these is the main task…

Computational Engineering, Finance, and Science · Computer Science 2024-12-11 Nicola Henkelmann , Stephan Rhode , Johannes von Keler

While machine translation evaluation metrics based on string overlap (e.g., BLEU) have their limitations, their computations are transparent: the BLEU score assigned to a particular candidate translation can be traced back to the presence…

Computation and Language · Computer Science 2022-10-26 Marzena Karpinska , Nishant Raj , Katherine Thai , Yixiao Song , Ankita Gupta , Mohit Iyyer

Radiology report generation aims to automatically generate a clinically accurate and coherent paragraph from the X-ray image, which could relieve radiologists from the heavy burden of report writing. Although various image caption methods…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Zhongzhen Huang , Xiaofan Zhang , Shaoting Zhang

We introduce a radiology-focused visual language model designed to generate radiology reports from chest X-rays. Building on previous findings that large language models (LLMs) can acquire multimodal capabilities when aligned with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xi Zhang , Zaiqiao Meng , Jake Lever , Edmond S. L. Ho

In response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic…

Image and Video Processing · Electrical Eng. & Systems 2024-05-24 Zhusi Zhong , Jie Li , John Sollee , Scott Collins , Harrison Bai , Paul Zhang , Terrence Healey , Michael Atalay , Xinbo Gao , Zhicheng Jiao

Automatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Qingqiu Li , Jilan Xu , Runtian Yuan , Mohan Chen , Yuejie Zhang , Rui Feng , Xiaobo Zhang , Shang Gao

Large language models (LLMs) like ChatGPT show excellent capabilities in various natural language processing tasks, especially for text generation. The effectiveness of LLMs in summarizing radiology report impressions remains unclear. In…

Computation and Language · Computer Science 2025-04-07 Danqing Hu , Shanyuan Zhang , Qing Liu , Xiaofeng Zhu , Bing Liu

Neural metrics for machine translation evaluation, such as COMET, exhibit significant improvements in their correlation with human judgments, as compared to traditional metrics based on lexical overlap, such as BLEU. Yet, neural metrics…

Computation and Language · Computer Science 2023-05-22 Ricardo Rei , Nuno M. Guerreiro , Marcos Treviso , Luisa Coheur , Alon Lavie , André F. T. Martins

Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed…

Computation and Language · Computer Science 2026-05-14 Yinzhu Chen , Abdine Maiga , Hossein A. Rahmani , Emine Yilmaz