English
Related papers

Related papers: Ran Score: a LLM-based Evaluation Score for Radiol…

200 papers

Advancing representation learning in specialized fields like medicine remains challenging due to the scarcity of expert annotations for text and images. To tackle this issue, we present a novel two-stage framework designed to extract…

Computation and Language · Computer Science 2024-07-03 Pablo Messina , René Vidal , Denis Parra , Álvaro Soto , Vladimir Araujo

Vision-language models (VLMs) are increasingly used for medical image interpretation, yet they frequently hallucinate, generating clinically plausible but factually incorrect findings that pose direct patient safety risks. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Haoyu Wang , Zitong Li

The chest X-rays (CXRs) is one of the views most commonly ordered by radiologists (NHS),which is critical for diagnosis of many different thoracic diseases. Accurately detecting thepresence of multiple diseases from CXRs is still a…

Computer Vision and Pattern Recognition · Computer Science 2020-05-27 Hieu H. Pham , Tung T. Le , Dat T. Ngo , Dat Q. Tran , Ha Q. Nguyen

Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenarios, clinicians often lack prior clinical knowledge, forcing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Junzhi Ning , Jiashi Lin , Yingying Fang , Wei Li , Jiyao Liu , Cheng Tang , Chenglong Ma , Wenhao Tang , Tianbin Li , Ziyan Huang , Guang Yang , Junjun He

The paper proposes a novel evaluation metric for automatic medical report generation from X-ray images, VLScore. It aims to overcome the limitations of existing evaluation methods, which either focus solely on textual similarities, ignoring…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Gefen Dawidowicz , Elad Hirsch , Ayellet Tal

Recently large vision-language models have shown potential when interpreting complex images and generating natural language descriptions using advanced reasoning. Medicine's inherently multimodal nature incorporating scans and text-based…

Image and Video Processing · Electrical Eng. & Systems 2024-07-15 Naman Sharma

Large language models (LLMs) show increasing promise in medical applications, but their ability to detect and correct errors in clinical texts -- a prerequisite for safe deployment -- remains under-evaluated, particularly beyond English. We…

Computation and Language · Computer Science 2025-11-04 Naoto Iwase , Hiroki Okuyama , Junichiro Iwasawa

Data labeling is currently a time-consuming task that often requires expert knowledge. In research settings, the availability of correctly labeled data is crucial to ensure that model predictions are accurate and useful. We propose…

Machine Learning · Computer Science 2018-12-31 Marina Bendersky , Joy Wu , Tanveer Syeda-Mahmood

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data…

Computation and Language · Computer Science 2026-03-26 Tianpeng Zheng , Zhehan Jiang , Jiayi Liu , Shicong Feng

Retrieval-augmented generation (RAG) systems are increasingly deployed in user-facing applications, yet systematic, human-centered evaluation of their outputs remains underexplored. Building on Gienapp's utility-dimension framework, we…

Artificial Intelligence · Computer Science 2025-10-01 Aline Mangold , Kiran Hoffmann

While human evaluation is the most reliable metric for evaluating speech generation systems, it is generally costly and time-consuming. Previous studies on automatic speech quality assessment address the problem by predicting human…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-12 Soumi Maiti , Yifan Peng , Takaaki Saeki , Shinji Watanabe

Radiology reports contain a diverse and rich set of clinical abnormalities documented by radiologists during their interpretation of the images. Comprehensive semantic representations of radiological findings would enable a wide range of…

Computation and Language · Computer Science 2021-12-28 Wilson Lau , Kevin Lybarger , Martin L. Gunn , Meliha Yetisgen

While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to…

Radiology report generation represents a significant application within medical AI, and has achieved impressive results. Concurrently, large language models (LLMs) have demonstrated remarkable performance across various domains. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Haifeng Zhao , Yufei Zhang , Leilei Ma , Shuo Xu , Dengdi Sun

Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Trong Thang Pham , Ngoc-Vuong Ho , Nhat-Tan Bui , Thinh Phan , Patel Brijesh , Donald Adjeroh , Gianfranco Doretto , Anh Nguyen , Carol C. Wu , Hien Nguyen , Ngan Le

Radiologists highly desire fully automated AI for radiology report generation (R2G), yet existing approaches fall short in clinical utility. Reinforcement learning (RL) holds potential to address these shortcomings, but its adoption in this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zilin Lu , Ruifeng Yuan , Weiwei Cao , Wanxing Chang , Zhongyu Wei , Sinuo Wang , Yong Xia , Ling Zhang , Jianpeng Zhang

Large Language Models (LLMs) excel in data synthesis but can be inaccurate in domain-specific tasks, which retrieval-augmented generation (RAG) systems address by leveraging user-provided data. However, RAGs require optimization in both…

Computation and Language · Computer Science 2024-11-05 Kazi Ahmed Asif Fuad , Lizhong Chen

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS),…

Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalanced model training, LLMs may achieve unreal evaluation results on public benchmarks, either…

Computation and Language · Computer Science 2026-05-26 Zi Liang , Liantong Yu , Shiyu Zhang , Qingqing Ye , Haibo Hu

Radiology report generation (RRG) aims to automatically produce diagnostic reports from medical images, with the potential to enhance clinical workflows and reduce radiologists' workload. While recent approaches leveraging multimodal large…

Artificial Intelligence · Computer Science 2025-05-16 Ziruo Yi , Ting Xiao , Mark V. Albert
‹ Prev 1 8 9 10 Next ›