English
Related papers

Related papers: GREEN: Generative Radiology Report Evaluation and …

200 papers

Radiology report generation aims at generating descriptive text from radiology images automatically, which may present an opportunity to improve radiology reporting and interpretation. A typical setting consists of training encoder-decoder…

Computation and Language · Computer Science 2021-09-28 An Yan , Zexue He , Xing Lu , Jiang Du , Eric Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

Evaluation metrics are a key ingredient for progress of text generation systems. In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, MoverScore, BLEURT, etc.) which correlate much better with…

Computation and Language · Computer Science 2021-11-02 Marvin Kaster , Wei Zhao , Steffen Eger

We present Head CT Ontology Normalized Evaluation (HeadCT-ONE), a metric for evaluating head CT report generation through ontology-normalized entity and relation extraction. HeadCT-ONE enhances current information extraction derived metrics…

Radiology reports play a critical role in communicating medical findings to physicians. In each report, the impression section summarizes essential radiology findings. In clinical practice, writing impression is highly demanded yet…

Computation and Language · Computer Science 2021-12-21 Jinpeng Hu , Jianling Li , Zhihong Chen , Yaling Shen , Yan Song , Xiang Wan , Tsung-Hui Chang

The integration of artificial intelligence (AI) into medical diagnostic workflows requires robust and consistent evaluation methods to ensure reliability, clinical relevance, and the inherent variability in expert judgments. Traditional…

Text generation is an important Natural Language Processing task with various applications. Although several metrics have already been introduced to evaluate the text generation methods, each of them has its own shortcomings. The most…

Machine Learning · Computer Science 2019-05-22 Ehsan Montahaei , Danial Alihosseini , Mahdieh Soleymani Baghshah

This study investigates how accurately different evaluation metrics capture the quality of causal explanations in automatically generated diagnostic reports. We compare six metrics: BERTScore, Cosine Similarity, BioSentVec, GPT-White,…

Computation and Language · Computer Science 2025-06-24 Yousang Cho , Key-Sun Choi

Radiology reporting is a complex task requiring detailed medical image understanding and precise language generation, for which generative multimodal models offer a promising solution. However, to impact clinical practice, models must…

Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses…

Recent advances in radiology report generation (RRG) have been driven by large paired image-text datasets; however, progress in neuro-oncology has been limited due to a lack of open paired image-report datasets. Here, we introduce BTReport,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Juampablo E. Heras Rivera , Dickson T. Chen , Tianyi Ren , Daniel K. Low , Asma Ben Abacha , Alberto Santamaria-Pang , Mehmet Kurt

Multimodal foundation models hold significant potential for automating radiology report generation, thereby assisting clinicians in diagnosing cardiac diseases. However, generated reports often suffer from serious factual inaccuracy. In…

Computation and Language · Computer Science 2025-02-07 Liwen Sun , James Zhao , Megan Han , Chenyan Xiong

The increasing prevalence of retinal diseases poses a significant challenge to the healthcare system, as the demand for ophthalmologists surpasses the available workforce. This imbalance creates a bottleneck in diagnosis and treatment,…

Image and Video Processing · Electrical Eng. & Systems 2024-08-15 Jia-Hong Huang

Medical images are widely used in clinical practice for diagnosis. Automatically generating interpretable medical reports can reduce radiologists' burden and facilitate timely care. However, most existing approaches to automatic report…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Jinghan Sun , Dong Wei , Liansheng Wang , Yefeng Zheng

The world faces a shortage of radiologists, leading to longer treatment times and increased stress, negatively impacting patient safety and workforce morale. Integrating artificial intelligence to interpret radiographic images and generate…

Image and Video Processing · Electrical Eng. & Systems 2024-06-19 Marijn Borghouts

In recent years, climate change repercussions have increasingly captured public interest. Consequently, corporations are emphasizing their environmental efforts in sustainability reports to bolster their public image. Yet, the absence of…

Computation and Language · Computer Science 2024-11-26 Avalon Vinella , Margaret Capetz , Rebecca Pattichis , Christina Chance , Reshmi Ghosh , Kai-Wei Chang

Automatic metrics are extensively used to evaluate natural language processing systems. However, there has been increasing focus on how they are used and reported by practitioners within the field. In this paper, we have conducted a survey…

Unlike classical lexical overlap metrics such as BLEU, most current evaluation metrics (such as BERTScore or MoverScore) are based on black-box language models such as BERT or XLM-R. They often achieve strong correlations with human…

Computation and Language · Computer Science 2022-03-22 Christoph Leiter , Piyawat Lertvittayakumjorn , Marina Fomicheva , Wei Zhao , Yang Gao , Steffen Eger

Automatic report generation has arisen as a significant research area in computer-aided diagnosis, aiming to alleviate the burden on clinicians by generating reports automatically based on medical images. In this work, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Jun Li , Tongkun Su , Baoliang Zhao , Faqin Lv , Qiong Wang , Nassir Navab , Ying Hu , Zhongliang Jiang

Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when applied to reports…

Artificial Intelligence · Computer Science 2026-04-07 Federica Bologna , Jean-Philippe Corbeil , Matthew Wilkens , Asma Ben Abacha

The vision-language modeling capability of multi-modal large language models has attracted wide attention from the community. However, in medical domain, radiology report generation using vision-language models still faces significant…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Yuhao Wang , Chao Hao , Yawen Cui , Xinqi Su , Weicheng Xie , Tao Tan , Zitong Yu