中文
相关论文

相关论文: Decomposing Physician Disagreement in HealthBench

200 篇论文

Objectives: To evaluate GPT-4o's ability to extract diagnostic labels (with uncertainty) from free-text radiology reports and to test how these labels affect multi-label image classification of musculoskeletal radiographs. Methods: This…

Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may…

The interpretation of chest radiographs is an essential task for the detection of thoracic diseases and abnormalities. However, it is a challenging problem with high inter-rater variability and inherent ambiguity due to inconclusive…

Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their…

计算与语言 · 计算机科学 2026-01-21 Kie Shidara , Preethi Prem , Jonathan Kim , Anna Podlasek , Feng Liu , Ahmed Alaa , Danilo Bernardo

The rapid growth of healthcare data and advances in computational power have accelerated the adoption of artificial intelligence (AI) in medicine. However, AI systems deployed without explicit fairness considerations risk exacerbating…

机器学习 · 计算机科学 2025-04-22 Xiaoyang Wang , Christopher C. Yang

Large language models (LLMs) show promise for extracting information from Electronic Health Records (EHR) and supporting clinical decisions. However, deployment in clinical settings faces challenges due to hallucination risks. We propose…

人工智能 · 计算机科学 2025-08-27 Yongwoo Song , Minbyul Jeong , Mujeen Sung

Medical image segmentation exhibits intra- and inter-annotator variability due to ambiguous object boundaries, annotator preferences, expertise, and tools, among other factors. Lesions with ambiguous boundaries, e.g., spiculated or…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Kumar Abhishek , Jeremy Kawahara , Ghassan Hamarneh

Obtaining datasets labeled to facilitate model development is a challenge for most machine learning tasks. The difficulty is heightened for medical imaging, where data itself is limited in accessibility and labeling requires costly time and…

We reframe clinician overrides of clinical AI recommendations as implicit preference data - the same signal structure exploited by reinforcement learning from human feedback (RLHF), but richer: the annotator is a domain expert, the…

机器学习 · 计算机科学 2026-05-19 Prabhjot Singh , Abhishek Gupta , Chris Betz , Abe Flansburg , Brett Ives , Sudeep Lama , Jung Hoon Son

As artificial intelligence (AI) systems approach and surpass expert human performance across a broad range of tasks, obtaining high-quality human supervision for evaluation and training becomes increasingly challenging. Our focus is on…

机器学习 · 计算机科学 2026-02-25 Ren Yin , Takashi Ishida , Masashi Sugiyama

Measuring cross-sectional areas in ultrasound images is a standard tool to evaluate disease progress or treatment response. Often addressed today with supervised deep-learning segmentation approaches, existing solutions highly depend upon…

图像与视频处理 · 电气工程与系统科学 2023-08-21 Vanessa Gonzalez Duque , Leonhard Zirus , Yordanka Velikova , Nassir Navab , Diana Mateus

One of the common issues in clinical decision-making is the presence of uncertainty, which often arises due to ambiguity in radiology reports, which often reflect genuine diagnostic uncertainty or limitations of automated label extraction…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Antik Aich Roy , Ujjwal Bhattacharya

Diagnostic performance in medical AI varies systematically across demographic groups, yet subgroup AUC can mask clinically important disparities. At a fixed inference-time operating point, some groups may exhibit over-diagnostic behaviour,…

Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bases -- trials, primary studies, regulatory documents, and…

Representation bias is one of the most common types of biases in artificial intelligence (AI) systems, causing AI models to perform poorly on underrepresented data segments. Although AI practitioners use various methods to reduce…

人机交互 · 计算机科学 2025-02-28 Aditya Bhattacharya , Simone Stumpf , Robin De Croon , Katrien Verbert

Medical large language models (LLMs) achieve impressive performance on standardized benchmarks, yet these evaluations fail to capture the complexity of real clinical encounters where patients exhibit memory gaps, limited health literacy,…

Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no standardized and…

人工智能 · 计算机科学 2024-12-18 Deep Bhatt , Surya Ayyagari , Anuruddh Mishra

Purpose: Artificial intelligence (AI) solutions for medical diagnosis require thorough evaluation to demonstrate that performance is maintained for all patient sub-groups and to ensure that proposed improvements in care will be delivered…

Background: Evaluating AI-generated treatment plans is a key challenge as AI expands beyond diagnostics, especially with new reasoning models. This study compares plans from human experts and two AI models (a generalist and a reasoner),…

人工智能 · 计算机科学 2025-07-09 Dipayan Sengupta , Saumya Panda

Artificial intelligence has shown significant promise in chest radiography, where deep learning models can approach radiologist-level diagnostic performance. Progress has been accelerated by large public datasets such as MIMIC-CXR,…

机器学习 · 计算机科学 2026-03-17 Amy Rafferty , Ajitha Rajan
‹ 上一页 1 8 9 10 下一页 ›