English
Related papers

Related papers: VERT: Reliable LLM Judges for Radiology Report Eva…

200 papers

As LLM-as-a-Judge emerges as a new paradigm for assessing large language models (LLMs), concerns have been raised regarding the alignment, bias, and stability of LLM evaluators. While substantial work has focused on alignment and bias,…

Computation and Language · Computer Science 2025-03-04 Qiujie Xie , Qingqiu Li , Zhuohao Yu , Yuejie Zhang , Yue Zhang , Linyi Yang

Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM jury composed of three frontier AI models scoring 3333…

Human relevance assessment is time-consuming and cognitively intensive, limiting the scalability of Information Retrieval evaluation. This has led to growing interest in using large language models (LLMs) as proxies for human judges.…

Information Retrieval · Computer Science 2026-04-28 Chuting Yu , Hang Li , Guido Zuccon , Joel Mackenzie , Teerapong Leelanupab

The integration of artificial intelligence in healthcare has opened new horizons for improving medical diagnostics and patient care. However, challenges persist in developing systems capable of generating accurate and contextually relevant…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Marco Salmè , Rosa Sicilia , Paolo Soda , Valerio Guarrasi

This study evaluates the performance of Large Language Models (LLMs) as an Artificial Intelligence-based tutor for a university course. In particular, different advanced techniques are utilized, such as prompt engineering,…

EXplainable machine learning (XML) has recently emerged to address the mystery mechanisms of machine learning (ML) systems by interpreting their 'black box' results. Despite the development of various explanation methods, determining the…

Human-Computer Interaction · Computer Science 2025-03-03 Bo Wang , Yiqiao Li , Jianlong Zhou , Fang Chen

Recent advancements in large language models have led to significant improvements across various tasks, including mathematical reasoning, which is used to assess models' intelligence in logical reasoning and problem-solving. Models are…

Artificial Intelligence · Computer Science 2026-04-27 Erez Yosef , Oron Anschel , Shunit Haviv Hakimi , Asaf Gendler , Adam Botach , Nimrod Berman , Igor Kviatkovsky

Interpreting quantitative CT biomarkers, such as organ volume and tissue attenuation, requires large-scale healthy reference distributions. However, creating these is challenging because clinical datasets are often heavily enriched with…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Christian Wachinger , Bernhard Renger , Christopher Späth , Jan Kirschke , Marcus Makowski

Large language models (LLMs) constitute a breakthrough state-of-the-art Artificial Intelligence technology which is rapidly evolving and promises to aid in medical diagnosis. However, the correctness and the accuracy of their returns has…

Computation and Language · Computer Science 2024-02-07 Dimitrios P. Panagoulias , Maria Virvou , George A. Tsihrintzis

Developing imaging models capable of detecting pathologies from chest X-rays can be cost and time-prohibitive for large datasets as it requires supervision to attain state-of-the-art performance. Instead, labels extracted from radiology…

Computation and Language · Computer Science 2024-08-09 Panagiotis Fytas , Anna Breger , Ian Selby , Simon Baker , Shahab Shahipasand , Anna Korhonen

The growing integration of vision-language models (VLMs) in medical applications offers promising support for diagnostic reasoning. However, current medical VLMs often face limitations in generalization, transparency, and computational…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Tan-Hanh Pham , Chris Ngo

PET/CT imaging is pivotal in oncology and nuclear medicine, yet summarizing complex findings into precise diagnostic impressions is labor-intensive. While LLMs have shown promise in medical text generation, their capability in the highly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Yuchen Liu , Wenbo Zhang , Liling Peng , Yichi Zhang , Yu Fu , Xin Guo , Chao Qu , Yuan Qi , Le Xue

This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields four key findings: 1) LRMs outperform non-reasoning LLMs in…

Computation and Language · Computer Science 2026-05-15 Hui Huang , Xuanxin Wu , Muyun Yang , Yuki Arase

Large language models (LLMs) like ChatGPT show excellent capabilities in various natural language processing tasks, especially for text generation. The effectiveness of LLMs in summarizing radiology report impressions remains unclear. In…

Computation and Language · Computer Science 2025-04-07 Danqing Hu , Shanyuan Zhang , Qing Liu , Xiaofeng Zhu , Bing Liu

General-purpose LLM judges capable of human-level evaluation provide not only a scalable and accurate way of evaluating instruction-following LLMs but also new avenues for supervising and improving their performance. One promising way of…

Computation and Language · Computer Science 2025-02-27 Ian Wu , Patrick Fernandes , Amanda Bertsch , Seungone Kim , Sina Pakazad , Graham Neubig

RAG systems are increasingly evaluated and optimized using LLM judges, an approach that is rapidly becoming the dominant paradigm for system assessment. Nugget-based approaches in particular are now embedded not only in evaluation…

Information Retrieval · Computer Science 2026-03-30 Laura Dietz , Bryan Li , Eugene Yang , Dawn Lawrie , William Walden , James Mayfield

Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust. This gap reveals fundamental flaws in how current metrics assess the quality of generated reports. We…

Computation and Language · Computer Science 2025-10-02 Ruochen Li , Jun Li , Bailiang Jian , Kun Yuan , Youxiang Zhu

We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics, from classic n-gram overlap (BLEU, ROUGE) and contextual measures (BERTScore) to clinical concept-based…

Radiology reports are often lengthy and unstructured, posing challenges for referring physicians to quickly identify critical imaging findings while increasing the risk of missed information. This retrospective study aimed to enhance…

Computation and Language · Computer Science 2025-06-05 Iryna Hartsock , Cyrillo Araujo , Les Folio , Ghulam Rasool

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more…

Artificial Intelligence · Computer Science 2025-04-08 Sijun Tan , Siyuan Zhuang , Kyle Montgomery , William Y. Tang , Alejandro Cuadron , Chenguang Wang , Raluca Ada Popa , Ion Stoica
‹ Prev 1 3 4 5 6 7 10 Next ›