中文
相关论文

相关论文: Decomposing Physician Disagreement in HealthBench

200 篇论文

Introduction: Healthcare AI models often inherit biases from their training data. While efforts have primarily targeted bias in structured data, mental health heavily depends on unstructured data. This study aims to detect and mitigate…

Although large language models (LLMs) demonstrate expert-level medical knowledge, aligning their open-ended outputs with fine-grained clinician preferences remains challenging. Existing methods often rely on coarse objectives or unreliable…

人工智能 · 计算机科学 2026-02-12 Shiwei Lyu , Xidong Wang , Lei Liu , Hao Zhu , Chaohe Zhang , Jian Wang , Jinjie Gu , Benyou Wang , Yue Shen

Sharing medical reports is essential for patient-centered care. A recent line of work has focused on automatically generating reports with NLP methods. However, different audiences have different purposes when writing/reading medical…

计算与语言 · 计算机科学 2023-05-16 Zexue He , An Yan , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior work has explored the ability of neural networks to quantify predictive, epistemic, and aleatoric…

机器学习 · 统计学 2025-08-07 Simon Baur , Wojciech Samek , Jackie Ma

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce \textsc{ClinConsensus}, a Chinese medical benchmark…

Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed…

计算与语言 · 计算机科学 2026-05-14 Yinzhu Chen , Abdine Maiga , Hossein A. Rahmani , Emine Yilmaz

Uncertainty is a fundamental challenge in medical practice, but current medical AI systems fail to explicitly quantify or communicate uncertainty in a way that aligns with clinical reasoning. Existing XAI works focus on interpreting model…

人工智能 · 计算机科学 2025-09-24 Xiuyi Fan

Systematic mislabelling affecting specific subgroups (i.e., label bias) in medical imaging datasets represents an understudied issue concerning the fairness of medical AI systems. In this work, we investigated how size and separability of…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Emma A. M. Stanley , Raghav Mehta , Mélanie Roschewitz , Nils D. Forkert , Ben Glocker

We document a fundamental paradox in AI transparency: explanations improve decisions when algorithms are correct but systematically worsen them when algorithms err. In an experiment with 257 medical students making 3,855 diagnostic…

综合经济学 · 经济学 2025-12-10 Manshu Khanna , Ziyi Wang , Lijia Wei , Lian Xue

Deploying deep learning (DL) models in medical applications relies on predictive performance and other critical factors, such as conveying trustworthy predictive uncertainty. Uncertainty estimation (UE) methods provide potential solutions…

图像与视频处理 · 电气工程与系统科学 2024-03-26 Kudaibergen Abutalip , Numan Saeed , Ikboljon Sobirov , Vincent Andrearczyk , Adrien Depeursinge , Mohammad Yaqub

Deep neural networks for medical image classification often fail to generalize consistently in clinical practice due to violations of the i.i.d. assumption and opaque decision-making. This paper examines interpretability in deep neural…

Artificial intelligence (AI) systems, particularly those based on deep learning models, have increasingly achieved expert-level performance in medical applications. However, there is growing concern that such AI systems may reflect and…

计算与语言 · 计算机科学 2025-04-25 Xiuying Chen , Tairan Wang , Juexiao Zhou , Zirui Song , Xin Gao , Xiangliang Zhang

Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks…

机器学习 · 计算机科学 2025-10-14 Christopher Chiu , Silviu Pitis , Mihaela van der Schaar

Large language models (LLMs) increasingly operate as autonomous agents that reason over external APIs to perform complex tasks. However, their reliability and agreement remain poorly characterized. We present a unified benchmarking…

信息检索 · 计算机科学 2026-04-28 Eyhab Al-Masri

Language models deployed in high-stakes professional settings face conflicting demands from users, institutional authorities, and professional norms. How models act when these demands conflict reveals a principal hierarchy -- an implicit…

人工智能 · 计算机科学 2026-05-13 Fangyi Yu , Nabeel Seedat , Jonathan Richard Schwarz , Andrew M. Bean

Benchmarking the capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth answers from experts. This ambiguity is not just limited to…

Patient education materials for solid-organ transplantation vary substantially across U.S. centers, yet no systematic method exists to quantify this heterogeneity at scale. We introduce a framework that grounds the same patient questions in…

信息检索 · 计算机科学 2026-03-24 Yubo Li , Ramayya Krishnan , Rema Padman

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert…

Extracting medical decisions from clinical notes is a key step for clinical decision support and patient-facing care summaries. We study how the linguistic characteristics of clinical decisions vary across decision categories and whether…

计算与语言 · 计算机科学 2026-02-05 Mohamed Elgaar , Hadi Amiri

The interpretation of medical images is a challenging task, often complicated by the presence of artifacts, occlusions, limited contrast and more. Most notable is the case of chest radiography, where there is a high inter-rater variability…