中文
相关论文

相关论文: Decomposing Physician Disagreement in HealthBench

200 篇论文

Biased information (recently termed bisinformation) continues to be taught in medical curricula, often long after having been debunked. In this paper, we introduce BRICC, a firstin-class initiative that seeks to mitigate medical…

As artificial intelligence systems move toward clinical deployment, ensuring reliable prediction behavior is fundamental for safety-critical decision-making tasks. One proposed safeguard is selective prediction, where models can defer…

机器学习 · 计算机科学 2026-05-25 L. Julián Lechuga López , Farah E. Shamout , Tim G. J. Rudner

As machine learning algorithms are more widely deployed in healthcare, the question of algorithmic fairness becomes more critical to examine. Our work seeks to identify and understand disparities in a deployed model that classifies…

计算机与社会 · 计算机科学 2020-12-15 Elisa Ferracane , Sandeep Konam

Despite the remarkable performance of supervised medical image segmentation models, relying on a large amount of labeled data is impractical in real-world situations. Semi-supervised learning approaches aim to alleviate this challenge using…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Yunyao Lu , Yihang Wu , Ahmad Chaddad , Tareef Daqqaq , Reem Kateb

Large-scale language models (LLMs) often offer clinical judgments based on incomplete information, increasing the risk of misdiagnosis. Existing studies have primarily evaluated confidence in single-turn, static settings, overlooking the…

计算与语言 · 计算机科学 2026-01-23 Zhiyao Ren , Yibing Zhan , Siyuan Liang , Guozheng Ma , Baosheng Yu , Dacheng Tao

Medical conversational AI (AI) plays a pivotal role in the development of safer and more effective medical dialogue systems. However, existing benchmarks and evaluation frameworks for assessing the information-gathering and diagnostic…

计算与语言 · 计算机科学 2026-01-08 Lecheng Gong , Weimin Fang , Ting Yang , Dongjie Tao , Chunxiao Guo , Peng Wei , Bo Xie , Jinqun Guan , Zixiao Chen , Fang Shi , Jinjie Gu , Junwei Liu

A key step in medical diagnosis is giving the patient a universally recognized label (e.g. Appendicitis) which essentially assigns the patient to a class(es) of patients with similar body failures. However, two patients having the same…

人工智能 · 计算机科学 2019-09-10 Moshe BenBassat

This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS, MODMA, and PDCH. First, we re-evaluate E-DAIC under strict subject-disjoint…

计算与语言 · 计算机科学 2026-05-26 Takehiro Ishikawa , Jon Duke

This paper presents a comprehensive analysis of how AI-generated medical responses are perceived and evaluated by non-experts. A total of 300 participants gave evaluations for medical responses that were either written by a medical doctor…

人机交互 · 计算机科学 2024-08-29 Shruthi Shekar , Pat Pataranutaporn , Chethan Sarabu , Guillermo A. Cecchi , Pattie Maes

Techniques for reliable rubric-based LLM evaluation -- ensemble judging, bias mitigation, few-shot calibration -- are scattered across papers with inconsistent terminology and partial implementations. We introduce Autorubric, an open-source…

计算与语言 · 计算机科学 2026-04-07 Delip Rao , Chris Callison-Burch

Large language models (LLMs) are increasingly deployed in clinical decision support, yet subtle demographic cues can influence their reasoning. Prior work has documented disparities in outputs across patient groups, but little is known…

计算与语言 · 计算机科学 2025-10-16 Rajarshi Ghosh , Abhay Gupta , Hudson McBride , Anurag Vaidya , Faisal Mahmood

Large language models are increasingly deployed to simulate patients for clinical training, research, and mental health tools, yet population-level validity remains largely untested. We introduce PsychBench, the first epidemiological audit…

计算机与社会 · 计算机科学 2026-04-21 Patrick Keough

Across healthcare, agentic artificial intelligence (AI) systems are increasingly promoted as capable of autonomous action, yet in practice they currently operate under near-total human oversight due to safety, regulatory, and liability…

Dermatological care via telemedicine often lacks the rich context of in-person visits. Clinicians must make diagnoses based on a handful of images and brief descriptions, without the benefit of physical exams, second opinions, or reference…

人工智能 · 计算机科学 2025-08-27 Karishma Thakrar , Shreyas Basavatia , Akshay Daftardar

Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single supervision signal. This paradigm overlooks the intrinsic uncertainty of low-agreement…

计算与语言 · 计算机科学 2026-05-14 Junyu Lu , Deyi Ji , Xuanyi Liu , Lanyun Zhu , Bo Xu , Liang Yang , Xian-Sheng Hua , Hongfei Lin

Clinical decisions are often required under incomplete information. Clinical experts must identify whether available information is sufficient for judgment, as both premature conclusion and unnecessary abstention can compromise patient…

人工智能 · 计算机科学 2026-02-27 Yusuke Watanabe , Yohei Kobashi , Takeshi Kojima , Yusuke Iwasawa , Yasushi Okuno , Yutaka Matsuo

While large language models are capable diagnostic tools, the impact of multi-agent topology on diagnostic accuracy remains underexplored. This study evaluates four agent topologies, Control (single agent), Hierarchical, Adversarial, and…

多智能体系统 · 计算机科学 2026-03-10 Ahmed Almasoud

Clinical studies often require understanding elements of a patient's narrative that exist only in free text clinical notes. To transform notes into structured data for downstream use, these elements are commonly extracted and normalized to…

计算与语言 · 计算机科学 2020-08-03 Monica Agrawal , Chloe O'Connell , Yasmin Fatemi , Ariel Levy , David Sontag

The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-form, multi-step outputs. To address this gap, we introduce…

Classical noisy-label theory predicts that downstream performance under weak supervision is bounded above by the labeler's accuracy, implying a sharp crossover: once a gold-trained classifier matches the labeler, weak labels stop helping…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Bruce Changlong Xu , Jose James , Alexander Ryu