中文
相关论文

相关论文: VERT: Reliable LLM Judges for Radiology Report Eva…

200 篇论文

Accurate disease classification from radiology reports is essential for many applications. While supervised fine-tuning (SFT) of lightweight LLMs improves accuracy, it can degrade reasoning. We propose a two-stage approach: SFT on disease…

人工智能 · 计算机科学 2026-04-22 Yishu Wei , Yi Lin , Adam Flanders , George Shih , Yifan Peng

The evaluation bottleneck in recommendation systems has become particularly acute with the rise of Generative AI, where traditional metrics fall short of capturing nuanced quality dimensions that matter in specialized domains like legal…

计算与语言 · 计算机科学 2025-12-30 Anu Pradhan , Alexandra Ortan , Apurv Verma , Madhavan Seshadri

As large language models (LLMs) continue to advance, reliable evaluation methods are essential particularly for open-ended, instruction-following tasks. LLM-as-a-Judge enables automatic evaluation using LLMs as evaluators, but its…

计算与语言 · 计算机科学 2025-06-17 Yusuke Yamauchi , Taro Yano , Masafumi Oyamada

Vision-language models (VLMs) often produce chain-of-thought (CoT) explanations that sound plausible yet fail to reflect the underlying decision process, undermining trust in high-stakes clinical use. Existing evaluations rarely catch this…

Large Language Models (LLMs) and causal learning each hold strong potential for clinical decision making (CDM). However, their synergy remains poorly understood, largely due to the lack of systematic benchmarks evaluating their integration…

机器学习 · 计算机科学 2025-11-14 Linna Wang , Zhixuan You , Qihui Zhang , Jiunan Wen , Ji Shi , Yimin Chen , Yusen Wang , Fanqi Ding , Ziliang Feng , Li Lu

Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation. Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Wenjun Hou , Yi Cheng , Kaishuai Xu , Heng Li , Yan Hu , Wenjie Li , Jiang Liu

Background: The positive predictive value (PPV) of large language model (LLM)-based proofreading for radiology reports is limited due to the low error prevalence. Purpose: To assess whether a three-pass LLM framework enhances PPV and…

计算与语言 · 计算机科学 2025-06-26 Songsoo Kim , Seungtae Lee , See Young Lee , Joonho Kim , Keechan Kan , Dukyong Yoon

LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g., a 1-5 rating vs. a True/False label). Existing diagnoses of…

Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-Judge framework, which uses prompted LLMs to evaluate…

计算与语言 · 计算机科学 2026-04-03 Michael Krumdick , Charles Lovering , Varshini Reddy , Seth Ebner , Chris Tanner

In this study, we leverage LLM to enhance the semantic analysis and develop similarity metrics for texts, addressing the limitations of traditional unsupervised NLP metrics like ROUGE and BLEU. We develop a framework where LLMs such as…

计算与语言 · 计算机科学 2024-02-22 Shaochen Xu , Zihao Wu , Huaqin Zhao , Peng Shu , Zhengliang Liu , Wenxiong Liao , Sheng Li , Andrea Sikora , Tianming Liu , Xiang Li

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human…

计算与语言 · 计算机科学 2025-05-14 Andreas Stephan , Dawei Zhu , Matthias Aßenmacher , Xiaoyu Shen , Benjamin Roth

Background: Structured radiology reports remains underdeveloped due to labor-intensive structuring and narrative-style reporting. Deep learning, particularly large language models (LLMs) like GPT-3.5, offers promise in automating the…

计算与语言 · 计算机科学 2024-10-30 Hidetoshi Matsuo , Mizuho Nishio , Takaaki Matsunaga , Koji Fujimoto , Takamichi Murakami

Safe deployment of Large Vision-Language Models (LVLMs) in radiology report generation requires not only accurate predictions but also clinically interpretable indicators of when outputs should be thoroughly reviewed, enabling selective…

The integration of large language models (LLMs) into medical practice offers transformative potential, yet their real-world clinical applicability remains constrained by critical alignment issues: (1) a misalignment between static…

人工智能 · 计算机科学 2025-12-05 Yongnan Jin , Xurui Li , Feng Cao , Liucun Gao , Juanjuan Yao

Systematic reviews (SR), in which experts summarize and analyze evidence across individual studies to provide insights on a specialized topic, are a cornerstone for evidence-based clinical decision-making, research, and policy. Given the…

计算与语言 · 计算机科学 2025-05-30 Christopher Polzak , Alejandro Lozano , Min Woo Sun , James Burgess , Yuhui Zhang , Kevin Wu , Serena Yeung-Levy

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide…

Automatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to…

The diagnosis of pathological images is often limited by expert availability and regional disparities, highlighting the importance of automated diagnosis using Vision-Language Models (VLMs). Traditional multimodal models typically emphasize…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Jianyu Wu , Hao Yang , Xinhua Zeng , Guibing He , Zhiyu Chen , Zihui Li , Xiaochuan Zhang , Yangyang Ma , Run Fang , Yang Liu

Assessing the quality of scientific research is essential for scholarly communication, yet widely used approaches face limitations in scalability, subjectivity, and time delay. Recent advances in large language models (LLMs) offer new…

信息检索 · 计算机科学 2026-04-21 Mengjia Wu , Yi Zhang , Robin Haunschild , Lutz Bornmann

Large language models (LLMs) have shown strong results on a range of applications, including regression and scoring tasks. Typically, one obtains outputs from an LLM via autoregressive sampling from the model's output distribution. We show…

计算与语言 · 计算机科学 2024-11-04 Michal Lukasik , Harikrishna Narasimhan , Aditya Krishna Menon , Felix Yu , Sanjiv Kumar