中文
相关论文

相关论文: Decomposing Physician Disagreement in HealthBench

200 篇论文

In recent years, AI systems in the medical domain have advanced significantly. However, despite outperforming humans, they are rarely used in practice since it is often not clear how they make their decisions. Optimal explanation and…

人机交互 · 计算机科学 2026-05-06 Julia Wagner , Tim Schlippe

Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of human experts. A key challenge, however, is that AI reasoning frequently diverges from…

人工智能 · 计算机科学 2026-05-25 Belona Sonna , Alban Grastien

Large language models (LLMs) hold substantial promise for clinical decision support. However, their widespread adoption in medicine, particularly in healthcare, is hindered by their propensity to generate false or misleading outputs, known…

Ambient AI "scribe" systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed without careful review, and high-quality human reference transcripts are often unavailable for…

声音 · 计算机科学 2026-04-17 Abdolamir Karbalaie , Fernando Seoane , Farhad Abtahi

AI-generated health misinformation poses unprecedented threats to patient safety and healthcare system trust globally. This white paper presents an explainable AI framework developed through the EPSRC INDICATE project to combat medical…

信息检索 · 计算机科学 2025-09-05 Sueun Hong , Shuojie Fu , Ovidiu Serban , Brianna Bao , James Kinross , Francesa Toni , Guy Martin , Uddhav Vaghela

This study investigates uncertainty quantification in large language models (LLMs) for medical applications, emphasizing both technical innovations and philosophical implications. As LLMs become integral to clinical decision-making,…

Medical tasks are prone to inter-rater variability due to multiple factors such as image quality, professional experience and training, or guideline clarity. Training deep learning networks with annotations from multiple raters is a common…

图像与视频处理 · 电气工程与系统科学 2023-01-13 Andreanne Lemay , Charley Gros , Enamundram Naga Karthik , Julien Cohen-Adad

Large language models (LLMs) are entering clinician workflows, yet evaluations rarely measure how clinician reasoning shapes model behavior during clinical interactions. We combined 61 New England Journal of Medicine Case Records with 92…

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing…

人工智能 · 计算机科学 2026-04-14 Ruiyang Li , Fang Liu , Licheng Jiao , Xinglin Xie , Jiayao Hao , Shuo Li , Xu Liu , Jingyi Yang , Lingling Li , Puhua Chen , Wenping Ma

While multiple studies have explored the relation between inter-rater variability and deep learning model uncertainty in medical segmentation tasks, little is known about the impact of individual rater style. This study quantifies rater…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Olivier Vincent , Charley Gros , Julien Cohen-Adad

Automated grading systems can efficiently score short-answer responses, yet they often fail to indicate when a grading decision is uncertain or potentially contentious. We introduce semantic entropy, a measure of variability across multiple…

人工智能 · 计算机科学 2025-08-07 Karrtik Iyer , Manikandan Ravikiran , Prasanna Pendse , Shayan Mohanty

Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biases unattributed and offering practitioners no actionable levers. We first establish that…

计算与语言 · 计算机科学 2026-05-28 Manan Uppadhyay , Prashant Kodali , Pranjal Chitale , Reshma Ramaprasad , Himanshu Beniwal , Sunayana Sitaram

Large language models (LLMs) are entering clinical workflows as decision support tools, yet how they respond to explicit patient value statements -- the core content of shared decision-making -- remains unmeasured. We conducted a factorial…

计算机与社会 · 计算机科学 2026-03-03 Sanjay Basu

Clinical document classification is essential for converting unstructured medical texts into standardised ICD-10 diagnoses, yet it faces challenges due to complex medical language, privacy constraints, and limited annotated datasets. Large…

计算与语言 · 计算机科学 2026-02-03 Akram Mustafa , Usman Naseem , Mostafa Rahimi Azghadi

Performance uncertainty quantification is essential for reliable validation and eventual clinical translation of medical imaging artificial intelligence (AI). Confidence intervals (CIs) play a central role in this process by indicating how…

Healthcare decision-making represents one of the most challenging domains for Artificial Intelligence (AI), requiring the integration of diverse knowledge sources, complex reasoning, and various external analytical tools. Current AI systems…

计算与语言 · 计算机科学 2025-06-03 Yexiao He , Ang Li , Boyi Liu , Zhewei Yao , Yuxiong He

Large language models (LLMs) have shown considerable potential in supporting medical diagnosis. However, their effective integration into clinical workflows is hindered by physicians' difficulties in perceiving and trusting LLM…

人机交互 · 计算机科学 2026-01-28 Yuansong Xu , Yichao Zhu , Haokai Wang , Yuchen Wu , Yang Ouyang , Hanlu Li , Wenzhe Zhou , Xinyu Liu , Chang Jiang , Quan Li

Reliable Large Language Models (LLMs) should abstain when confidence is insufficient. However, prior studies often treat refusal as a generic "I don't know'', failing to distinguish input-level ambiguity (data uncertainty) from capability…

计算与语言 · 计算机科学 2026-04-21 Jingyi Ren , Ante Wang , Yunghwei Lai , Xiaolong Wang , Linlu Gong , Weitao Li , Weizhi Ma , Yang Liu

Multi-agent deliberation systems using large language models (LLMs) are increasingly proposed for policy simulation, yet they suffer from artificial consensus: evaluator agents converge on the same option regardless of their assigned value…

多智能体系统 · 计算机科学 2026-04-30 Ariel Sela

Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for clinical trust. We present RECAP (Reflect-Extract-Calibrate-Align-Produce), an inference-time…

计算与语言 · 计算机科学 2026-05-05 Adarsh Srinivasan , Jacob Dineen , Muhammad Umar Afzal , Muhammad Uzair Sarfraz , Irbaz B. Riaz , Ben Zhou