English
Related papers

Related papers: RadEval: A framework for radiology text evaluation

200 papers

Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Chengrun Li , Corentin Royer , Haozhe Luo , Bastian Wittmann , Xia Li , Ibrahim Hamamci , Sezgin Er , Anjany Sekuboyina , Bjoern Menze

Multimodal Large Language Models (MLLMs) have facilitated Multimodal Summarization with Multimodal Output (MSMO), wherein systems generate concise textual summaries accompanied by salient visuals from multimodal sources. However, current…

Artificial Intelligence · Computer Science 2026-05-13 Abid Ali , Diego Molla-Aliod , Usman Naseem

Evaluating the quality of automatically generated image descriptions is a complex task that requires metrics capturing various dimensions, such as grammaticality, coverage, accuracy, and truthfulness. Although human evaluation provides…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Jia-Hong Huang , Hongyi Zhu , Yixian Shen , Stevan Rudinac , Evangelos Kanoulas

Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property: a model must recognise when the evidential basis for an answer has failed. We study…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Hanqi Jiang , Junhao Chen , Mingyu Kang , Hyeokjae Kwon , Yi Pan , Lifeng Chen , Weihang You , Haozhen Gong , Ruiyu Yan , Jinglei Lv , Lin Zhao , Hui Ren , Quanzheng Li , Tianming Liu , Xiang Li

Developing imaging models capable of detecting pathologies from chest X-rays can be cost and time-prohibitive for large datasets as it requires supervision to attain state-of-the-art performance. Instead, labels extracted from radiology…

Computation and Language · Computer Science 2024-08-09 Panagiotis Fytas , Anna Breger , Ian Selby , Simon Baker , Shahab Shahipasand , Anna Korhonen

Recent advancements in artificial intelligence have significantly improved the automatic generation of radiology reports. However, existing evaluation methods fail to reveal the models' understanding of radiological images and their…

Artificial Intelligence · Computer Science 2024-08-27 Xiaoman Zhang , Julián N. Acosta , Hong-Yu Zhou , Pranav Rajpurkar

Despite the routine use of electronic health record (EHR) data by radiologists to contextualize clinical history and inform image interpretation, the majority of deep learning architectures for medical imaging are unimodal, i.e., they only…

Image and Video Processing · Electrical Eng. & Systems 2021-11-30 Yuyin Zhou , Shih-Cheng Huang , Jason Alan Fries , Alaa Youssef , Timothy J. Amrhein , Marcello Chang , Imon Banerjee , Daniel Rubin , Lei Xing , Nigam Shah , Matthew P. Lungren

Mastering medical knowledge is crucial for medical-specific LLMs. However, despite the existence of medical benchmarks like MedQA, a unified framework that fully leverages existing knowledge bases to evaluate LLMs' mastery of medical…

Computation and Language · Computer Science 2024-10-03 Yuxuan Zhou , Xien Liu , Chen Ning , Xiao Zhang , Ji Wu

Advancing representation learning in specialized fields like medicine remains challenging due to the scarcity of expert annotations for text and images. To tackle this issue, we present a novel two-stage framework designed to extract…

Computation and Language · Computer Science 2024-07-03 Pablo Messina , René Vidal , Denis Parra , Álvaro Soto , Vladimir Araujo

Radiology reports are a rich resource for advancing deep learning applications in medicine by leveraging the large volume of data continuously being updated, integrated, and shared. However, there are significant challenges as well, largely…

Information Retrieval · Computer Science 2017-11-21 Imon Banerjee , Sriraman Madhavan , Roger Eric Goldman , Daniel L. Rubin

Learned representations of scientific documents can serve as valuable input features for downstream tasks without further fine-tuning. However, existing benchmarks for evaluating these representations fail to capture the diversity of…

Computation and Language · Computer Science 2023-11-14 Amanpreet Singh , Mike D'Arcy , Arman Cohan , Doug Downey , Sergey Feldman

Radiology reporting is a crucial part of the communication between radiologists and other medical professionals, but it can be time-consuming and error-prone. One approach to alleviate this is structured reporting, which saves time and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Chantal Pellegrini , Matthias Keicher , Ege Özsoy , Nassir Navab

Automatic radiology report summarization is a crucial clinical task, whose key challenge is to maintain factual accuracy between produced summaries and ground truth radiology findings. Existing research adopts reinforcement learning to…

Computation and Language · Computer Science 2023-03-17 Qianqian Xie , Jiayu Zhou , Yifan Peng , Fei Wang

Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semantic-spatial gap between free-text narratives and volumetric anatomy. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Shuo Jiang , Yuhao Hong , Chunbo Jiang , Weihong Chen , Huangwei Chen , Shenghao Zhu , Beining Wu , Mingxuan Liu , Zhu Zhu , Feiwei Qin , Min Tan , Yifei Chen

Robust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems. However, current LLM-based evaluation frameworks predominantly rely on directly prompting resource-intensive models with complex…

Computation and Language · Computer Science 2025-05-29 Kun Li , Yunxiang Li , Tianhua Zhang , Hongyin Luo , Xixin Wu , James Glass , Helen Meng

Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft responses. However, LLM outputs may contain clinical inaccuracies,…

Computation and Language · Computer Science 2025-09-29 Wenyuan Chen , Fateme Nateghi Haredasht , Kameron C. Black , Francois Grolleau , Emily Alsentzer , Jonathan H. Chen , Stephen P. Ma

State estimation is an essential component of autonomous systems, usually relying on sensor fusion that integrates data from cameras, LiDARs and IMUs. Recently, radars have shown the potential to improve the accuracy and robustness of state…

Robotics · Computer Science 2024-06-28 Vlaho-Josip Štironja , Luka Petrović , Juraj Peršić , Ivan Marković , Ivan Petrović

The reliable evaluation of large language models (LLMs) in medical applications remains an open challenge, particularly in capturing the complexity of multi-turn doctor-patient interactions that unfold in real clinical environments.…

Artificial Intelligence · Computer Science 2025-10-15 Yuechun Yu , Han Ying , Haoan Jin , Wenjian Jiang , Dong Xian , Binghao Wang , Zhou Yang , Mengyue Wu

Large Language Models (LLMs) drive scientific question-answering on modern search engines, yet their evaluation robustness remains underexplored. We introduce YESciEval, an open-source framework that combines fine-grained rubric-based…

Computation and Language · Computer Science 2025-05-30 Jennifer D'Souza , Hamed Babaei Giglou , Quentin Münch

Recent advancements in multimodal models have significantly improved vision-language (VL) alignment in radiology. However, existing approaches struggle to effectively utilize complex radiology reports for learning and offer limited…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Jonggwon Park , Byungmu Yoon , Soobum Kim , Kyoyun Choi