中文
相关论文

相关论文: Decomposing Physician Disagreement in HealthBench

200 篇论文

Classifiers for medical image analysis are often trained with a single consensus label, based on combining labels given by experts or crowds. However, disagreement between annotators may be informative, and thus removing it may not be the…

计算机视觉与模式识别 · 计算机科学 2018-08-20 Veronika Cheplygina , Josien P. W. Pluim

Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited and lacks standardized evaluation frameworks. Such editing could revolutionize clinical practices by…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Minghao Liu , Zhitao He , Zhiyuan Fan , Qingyun Wang , Yi R. Fung

Deep neural networks excel in medical imaging but remain prone to biases, leading to fairness gaps across demographic groups. We provide the first systematic exploration of Human-AI alignment and fairness in this domain. Our results show…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Haozhe Luo , Ziyu Zhou , Zixin Shu , Aurélie Pahud de Mortanges , Robert Berke , Mauricio Reyes

Radiology reports are invaluable for clinical decision-making and hold great potential for automated analysis when structured into machine-readable formats. These reports often contain uncertainty, which we categorize into two distinct…

计算与语言 · 计算机科学 2026-03-02 Paloma Rabaey , Jong Hak Moon , Jung-Oh Lee , Min Gwan Kim , Hangyul Yoon , Thomas Demeester , Edward Choi

Inaccuracies in existing or generated clinical text may lead to serious adverse consequences, especially if it is a misdiagnosis or incorrect treatment suggestion. With Large Language Models (LLMs) increasingly being used across diverse…

When encountering a dubious diagnostic case, medical instance retrieval can help radiologists make evidence-based diagnoses by finding images containing instances similar to a query case from a large image database. The similarity between…

信息检索 · 计算机科学 2021-05-20 Jiansheng Fang , Huazhu Fu , Dan Zeng , Xiao Yan , Yuguang Yan , Jiang Liu

Medicine and deep learning-based artificial intelligence (AI) engineering represent two distinct fields each with decades of published history. With such history comes a set of terminology that has a specific way in which it is applied.…

Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce…

计算与语言 · 计算机科学 2026-01-23 Sravanthi Machcha , Sushrita Yerra , Sahil Gupta , Aishwarya Sahoo , Sharmin Sultana , Hong Yu , Zonghai Yao

Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for Area Chairs and editors to reliably identify and interpret such disagreements. Existing…

计算与语言 · 计算机科学 2026-05-12 Sandeep Kumar , Yash Kamdar , Abid Hossain , Bharti Kumari , Tanik Saikh , Asif Ekbal

Large language models are increasingly integrated into decision-making in areas such as healthcare, law, finance, engineering, and government. Yet they share a critical limitation: they produce fluent outputs even when their internal…

人工智能 · 计算机科学 2026-04-17 Rikard Rosenbacke , Carl Rosenbacke , Victor Rosenbacke , Martin McKee

We introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two…

人工智能 · 计算机科学 2025-06-09 Yuxin Zuo , Shang Qu , Yifei Li , Zhangren Chen , Xuekai Zhu , Ermo Hua , Kaiyan Zhang , Ning Ding , Bowen Zhou

Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall. We address this by separating 11…

This paper contributes with a pragmatic evaluation framework for explainable Machine Learning (ML) models for clinical decision support. The study revealed a more nuanced role for ML explanation models, when these are pragmatically embedded…

Automated short-answer grading (ASAG) remains a challenging task due to the linguistic variability of student responses and the need for nuanced, rubric-aligned partial credit. While Large Language Models (LLMs) offer a promising solution,…

计算与语言 · 计算机科学 2026-01-15 Haotian Deng , Chris Farber , Jiyoon Lee , David Tang

Early detection and rapid intervention of lung cancer are crucial. Nonetheless, ensuring an accurate diagnosis is challenging, as physicians' ability to interpret chest X-rays varies significantly depending on their experience and degree of…

图像与视频处理 · 电气工程与系统科学 2025-08-20 Hyeonjin Choi , Jinse Kim , Dong-yeon Yoo , Ju-sung Sun , Jung-won Lee

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasoning capability either…

Multi-agent LLM systems, where multiple prompted instances of a language model independently answer questions, are increasingly used for complex reasoning tasks. However, existing methods for quantifying the uncertainty of their collective…

计算与语言 · 计算机科学 2026-03-24 Bo Jiang

As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in audio. We evaluated these models on 170 clinical cases, each…

计算与语言 · 计算机科学 2025-11-11 Zhi Rui Tam , Yun-Nung Chen

Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial integrity. We present an automated system for detecting dosing errors in unstructured…

人工智能 · 计算机科学 2026-04-23 Mohammad AL-Smadi

Accurate medical image segmentation is crucial for diagnosis and analysis. However, the models without calibrated uncertainty estimates might lead to errors in downstream analysis and exhibit low levels of robustness. Estimating the…

图像与视频处理 · 电气工程与系统科学 2021-09-16 Yanwu Yang , Xutao Guo , Yiwei Pan , Pengcheng Shi , Haiyan Lv , Ting Ma