English
Related papers

Related papers: RadEval: A framework for radiology text evaluation

200 papers

Evaluating generative image models remains a difficult problem. This is due to the high dimensionality of the outputs, the challenging task of representing but not replicating training data, and the lack of metrics that fully correspond to…

Human-Computer Interaction · Computer Science 2023-01-12 Yannick Assogba , Adam Pearce , Madison Elliott

Radiology reports convey detailed clinical observations and capture diagnostic reasoning that evolves over time. However, existing evaluation methods are limited to single-report settings and rely on coarse metrics that fail to capture…

Evaluating tables qualitatively and quantitatively poses a significant challenge, as standard metrics often overlook subtle structural and content-level discrepancies. To address this, we propose a rubric-based evaluation framework that…

Computation and Language · Computer Science 2026-04-22 Vihang Pancholi , Jainit Bafna , Tejas Anvekar , Manish Shrivastava , Vivek Gupta

This study presents three deidentified large medical text datasets, named DISCHARGE, ECHO and RADIOLOGY, which contain 50K, 16K and 378K pairs of report and summary that are derived from MIMIC-III, respectively. We implement convincing…

Computation and Language · Computer Science 2023-02-09 Yunqi Zhu , Xuebing Yang , Yuanyuan Wu , Wensheng Zhang

Evaluating long-context radiology report generation is challenging. NLG metrics fail to capture clinical correctness, while LLM-based metrics often lack generalizability. Clinical accuracy metrics are more relevant but are sensitive to…

Computation and Language · Computer Science 2025-05-26 Ibrahim Ethem Hamamci , Sezgin Er , Suprosanna Shit , Hadrien Reynaud , Bernhard Kainz , Bjoern Menze

Despite impressive visual fidelity, current text-to-image (T2I) diffusion models struggle to depict rare, complex, or culturally nuanced concepts due to training data limitations. We introduce RAVEL, a training-free framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Kavana Venkatesh , Yusuf Dalva , Ismini Lourentzou , Pinar Yanardag

An infrastructure for multisite, geographically-distributed creation and collection of diverse, high-quality, curated and labeled radiology image data is crucial for the successful automated development, deployment, monitoring and…

Machine Learning · Computer Science 2020-09-01 Menashe Benjamin , Guy Engelhard , Alex Aisen , Yinon Aradi , Elad Benjamin

We address a fundamental challenge in Natural Language Generation (NLG) model evaluation -- the design and evaluation of evaluation metrics. Recognizing the limitations of existing automatic metrics and noises from how current human…

Computation and Language · Computer Science 2023-10-24 Ziang Xiao , Susu Zhang , Vivian Lai , Q. Vera Liao

Event cameras are a new type of vision sensor that incorporates asynchronous and independent pixels, offering advantages over traditional frame-based cameras such as high dynamic range and minimal motion blur. However, their output is not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Burak Ercan , Onur Eker , Aykut Erdem , Erkut Erdem

Evaluating text revision in scientific writing remains a challenge, as traditional metrics such as ROUGE and BERTScore primarily focus on similarity rather than capturing meaningful improvements. In this work, we analyse and identify the…

Computation and Language · Computer Science 2026-01-26 Léane Jourdan , Florian Boudin , Richard Dufour , Nicolas Hernandez

Medical text embedding models are foundational to a wide array of healthcare applications, ranging from clinical decision support and biomedical information retrieval to medical question answering, yet they remain hampered by two critical…

Computation and Language · Computer Science 2025-08-07 Mohammad Khodadad , Ali Shiraee Kasmaee , Mahdi Astaraki , Hamidreza Mahyar

Radiology report generation is critical for efficiency but current models lack the structured reasoning of experts, hindering clinical trust and explainability by failing to link visual findings to precise anatomical locations. This paper…

Artificial Intelligence · Computer Science 2026-03-03 Peiyuan Jing , Kinhei Lee , Zhenxuan Zhang , Huichi Zhou , Zhengqing Yuan , Zhifan Gao , Lei Zhu , Giorgos Papanastasiou , Yingying Fang , Guang Yang

Vision-language models (VLMs) have shown impressive abilities across a range of multi-modal tasks. However, existing metrics for evaluating the quality of text generated by VLMs typically focus on an overall evaluation for a specific task,…

Computation and Language · Computer Science 2026-03-10 Masanari Ohi , Masahiro Kaneko , Naoaki Okazaki , Nakamasa Inoue

We present MGTEVAL, an extensible platform for systematic evaluation of Machine-Generated Text (MGT) detectors. Despite rapid progress in MGT detection, existing evaluations are often fragmented across datasets, preprocessing, attacks, and…

Cryptography and Security · Computer Science 2026-04-29 Yuanfan Li , Qi Zhou , Chengzhengxu Li , Zhaohan Zhang , Chenxu Zhao , Zepu Ruan , Chao Shen , Xiaoming Liu

Existing mainstream approaches follow the encoder-decoder paradigm for generating radiology reports. They focus on improving the network structure of encoders and decoders, which leads to two shortcomings: overlooking the modality gap and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Yuanjiang Luo , Hongxiang Li , Xuan Wu , Meng Cao , Xiaoshuang Huang , Zhihong Zhu , Peixi Liao , Hu Chen , Yi Zhang

Medical imaging plays a pivotal role in diagnosis and treatment in clinical practice. Inspired by the significant progress in automatic image captioning, various deep learning (DL)-based methods have been proposed to generate radiology…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Yixin Wang , Zihao Lin , Zhe Xu , Haoyu Dong , Jiang Tian , Jie Luo , Zhongchao Shi , Yang Zhang , Jianping Fan , Zhiqiang He

We propose and demonstrate a novel machine learning algorithm that assesses pulmonary edema severity from chest radiographs. While large publicly available datasets of chest radiographs and free-text radiology reports exist, only limited…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Geeticka Chauhan , Ruizhi Liao , William Wells , Jacob Andreas , Xin Wang , Seth Berkowitz , Steven Horng , Peter Szolovits , Polina Golland

As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications,…

Computation and Language · Computer Science 2026-01-12 Jean-Philippe Corbeil , Minseon Kim , Maxime Griot , Sheela Agarwal , Alessandro Sordoni , Francois Beaulieu , Paul Vozila

We present an automated way to evaluate the text alignment of text-to-image generative diffusion models using standard image-text recognition datasets. Our method, called SelfEval, uses the generative model to compute the likelihood of real…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Sai Saketh Rambhatla , Ishan Misra

Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify appropriate benchmarks, reproduce heterogeneous evaluation…

Computation and Language · Computer Science 2026-03-11 Chengyu Shen , Yanheng Hou , Minghui Pan , Runming He , Zhen Hao Wong , Meiyi Qiang , Zhou Liu , Hao Liang , Peichao Lai , Zeang Sheng , Wentao Zhang