English
Related papers

Related papers: Evaluating AI systems under uncertain ground truth…

200 papers

Labelled "ground truth" datasets are routinely used to evaluate and audit AI algorithms applied in high-stakes settings. However, there do not exist widely accepted benchmarks for the quality of labels in these datasets. We provide…

Computation and Language · Computer Science 2021-11-18 Abhilash Mishra , Yash Gorana

Despite the growing promise of artificial intelligence (AI) in supporting decision-making across domains, fostering appropriate human reliance on AI remains a critical challenge. In this paper, we investigate the utility of exploring…

Human-Computer Interaction · Computer Science 2025-05-26 Min Hun Lee , Martyn Zhe Yu Tok

Despite the recent improvements in overall accuracy, deep learning systems still exhibit low levels of robustness. Detecting possible failures is critical for a successful clinical integration of these systems, where each data point…

Image and Video Processing · Electrical Eng. & Systems 2019-10-14 Alain Jungo , Mauricio Reyes

This study investigates uncertainty quantification in large language models (LLMs) for medical applications, emphasizing both technical innovations and philosophical implications. As LLMs become integral to clinical decision-making,…

Artificial Intelligence · Computer Science 2025-04-08 Zahra Atf , Seyed Amir Ahmad Safavi-Naini , Peter R. Lewis , Aref Mahjoubfar , Nariman Naderi , Thomas R. Savage , Ali Soroush

Background: Clinical prediction models for a health condition are commonly evaluated regarding performance for a population, although decisions are made for individuals. The classic view relates uncertainty in risk estimates for individuals…

AI systems have the potential to improve decision-making, but decision makers face the risk that the AI may be misaligned with their objectives. We study this problem in the context of a treatment decision, where a designer decides which…

Theoretical Economics · Economics 2025-09-19 Drew Fudenberg , Annie Liang

Bias in medical artificial intelligence is conventionally viewed as a defect requiring elimination. However, human reasoning inherently incorporates biases shaped by education, culture, and experience, suggesting their presence may be…

Artificial Intelligence · Computer Science 2026-03-05 Farhad Abtahi , Mehdi Astaraki , Fernando Seoane

Estimating the test performance of software AI-based medical devices under distribution shifts is crucial for evaluating the safety, efficiency, and usability prior to clinical deployment. Due to the nature of regulated medical device…

Machine Learning · Computer Science 2022-07-14 Charles Lu , Syed Rakin Ahmed , Praveer Singh , Jayashree Kalpathy-Cramer

This paper aims to evaluate the suitability of current deep learning methods for clinical workflow especially by focusing on dermatology. Although deep learning methods have been attempted to get dermatologist level accuracy in several…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Sourav Mishra , Subhajit Chaudhury , Hideaki Imaizumi , Toshihiko Yamasaki

Reliable uncertainty quantification (UQ) is essential in medical AI. Evidential Deep Learning (EDL) offers a computationally efficient way to quantify model uncertainty alongside predictions, unlike traditional methods such as Monte Carlo…

There are a lot of hidden dangers in the change of human skin conditions, such as the sunburn caused by long-time exposure to ultraviolet radiation, which not only has aesthetic impact causing psychological depression and lack of…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Min Chen , Ping Zhou , Di Wu , Long Hu , Mohammad Mehedi Hassan , Atif Alamri

Estimating and disentangling epistemic uncertainty, uncertainty that is reducible with more training data, and aleatoric uncertainty, uncertainty that is inherent to the task at hand, is critically important when applying machine learning…

Machine Learning · Computer Science 2024-11-08 Matthew A. Chan , Maria J. Molina , Christopher A. Metzler

We consider a patient risk models which has access to patient features such as vital signs, lab values, and prior history but does not have access to a patient's diagnosis. For example, this occurs in a model deployed at intake time for…

Artificial Intelligence · Computer Science 2023-07-03 Alexander Peysakhovich , Rich Caruana , Yin Aphinyanaphongs

AI and ML models have already found many applications in critical domains, such as healthcare and criminal justice. However, fully automating such high-stakes applications can raise ethical or fairness concerns. Instead, in such cases,…

Artificial Intelligence · Computer Science 2023-04-28 Ioannis Papantonis , Vaishak Belle

AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can…

Human-Computer Interaction · Computer Science 2025-08-12 Sarah Jabbour , David Fouhey , Nikola Banovic , Stephanie D. Shepard , Ella Kazerooni , Michael W. Sjoding , Jenna Wiens

A growing literature on human-AI decision-making investigates strategies for combining human judgment with statistical models to improve decision-making. Research in this area often evaluates proposed improvements to models, interfaces, or…

Computers and Society · Computer Science 2023-05-29 Luke Guerdan , Amanda Coston , Zhiwei Steven Wu , Kenneth Holstein

This paper examines what it means for a medical AI system to be right by grounding the question in a specific clinical context: the automatic classification of plasma cells in digitized bone marrow smears for the diagnosis of multiple…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Antony Gitau

In medical imaging, inter-observer variability among radiologists often introduces label uncertainty, particularly in modalities where visual interpretation is subjective. Lung ultrasound (LUS) is a prime example-it frequently presents a…

Artificial intelligence (AI) systems increasingly achieve expert-level predictive accuracy in healthcare, yet improvements in model performance often fail to produce corresponding gains in patient outcomes. We term this disconnect the…

Artificial Intelligence · Computer Science 2026-01-13 Rifa Ferzana