English
Related papers

Related papers: Evaluating AI systems under uncertain ground truth…

200 papers

Machine learning risks reinforcing biases present in data and, as we argue in this work, in what is absent from data. In healthcare, societal and decision biases shape patterns in missing data, yet the algorithmic fairness implications of…

Artificial Intelligence · Computer Science 2025-03-19 Vincent Jeanselme , Maria De-Arteaga , Zhe Zhang , Jessica Barrett , Brian Tom

The ability to acknowledge the inevitable uncertainty in their knowledge and reasoning is a prerequisite for AI systems to be truly truthful and reliable. In this paper, we present a taxonomy of uncertainty specific to vision-language AI…

Artificial Intelligence · Computer Science 2024-07-03 Khyathi Raghavi Chandu , Linjie Li , Anas Awadalla , Ximing Lu , Jae Sung Park , Jack Hessel , Lijuan Wang , Yejin Choi

We evaluate artificial intelligence (AI) systems without ground truth by exploiting a link between strategic gaming and information loss. Building on established information theory, we analyze which mechanisms resist adversarial…

Machine Learning · Computer Science 2026-05-01 Zachary Robertson , Sanmi Koyejo

Reliable evaluation of AI systems remains a fundamental challenge when ground truth labels are unavailable, particularly for systems generating natural language outputs like AI chat and agent systems. Many of these AI agents and systems…

Machine Learning · Statistics 2025-11-05 Kaihua Ding

The deployment of AI systems for welfare benefit allocation allows for accelerated decision-making and faster provision of critical help, but has already led to an increase in unfair benefit denials and false fraud accusations. Collecting…

Computers and Society · Computer Science 2024-07-18 Mengchen Dong , Jean-François Bonnefon , Iyad Rahwan

Early detection and rapid intervention of lung cancer are crucial. Nonetheless, ensuring an accurate diagnosis is challenging, as physicians' ability to interpret chest X-rays varies significantly depending on their experience and degree of…

Image and Video Processing · Electrical Eng. & Systems 2025-08-20 Hyeonjin Choi , Jinse Kim , Dong-yeon Yoo , Ju-sung Sun , Jung-won Lee

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct…

Computation and Language · Computer Science 2026-05-27 Shang Wu , Randol Yao

The assessment of process mining techniques using real-life data is often compromised by the lack of ground truth knowledge, the presence of non-essential outliers in system behavior and recording errors in event logs. Using synthetically…

Databases · Computer Science 2025-01-27 Dominique Sommers , Natalia Sidorova , Boudewijn van Dongen

In safety-critical applications like medical diagnosis, certainty associated with a model's prediction is just as important as its accuracy. Consequently, uncertainty estimation and reduction play a crucial role. Uncertainty in predictions…

Image and Video Processing · Electrical Eng. & Systems 2023-09-12 Abhishek Singh Sambyal , Narayanan C. Krishnan , Deepti R. Bathula

Although machine learning (ML) models of AI achieve high performances in medicine, they are not free of errors. Empowering clinicians to identify incorrect model recommendations is crucial for engendering trust in medical AI. Explainable AI…

Artificial Intelligence · Computer Science 2022-12-20 Isil Guzey , Ozlem Ucar , Nukhet Aladag Ciftdemir , Betul Acunas

Medicine and deep learning-based artificial intelligence (AI) engineering represent two distinct fields each with decades of published history. With such history comes a set of terminology that has a specific way in which it is applied.…

Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system identifies, how it prioritizes them, or whether those priorities align with the review…

Artificial Intelligence · Computer Science 2026-04-23 Ming Jin

Artificial intelligence (AI) has demonstrated the ability to extract insights from data, but the issue of fairness remains a concern in high-stakes fields such as healthcare. Despite extensive discussion and efforts in algorithm…

Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels).…

Machine Learning · Computer Science 2021-10-29 Maria Ulan , Welf Löwe , Morgan Ericsson , Anna Wingkvist

The demand for Explainable AI (XAI) has triggered an explosion of methods, producing a landscape so fragmented that we now rely on surveys of surveys. Yet, fundamental challenges persist: conflicting metrics, failed sanity checks, and…

Machine Learning · Computer Science 2026-03-31 Amir-Hossein Karimi

Deep learning algorithms have demonstrated remarkable efficacy in various medical image analysis (MedIA) applications. However, recent research highlights a performance disparity in these algorithms when applied to specific subgroups, such…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Zikang Xu , Jun Li , Qingsong Yao , Han Li , Mingyue Zhao , S. Kevin Zhou

Machine learning (ML) models may suffer from significant performance disparities between patient groups. Identifying such disparities by monitoring performance at a granular level is crucial for safely deploying ML to each patient.…

Machine Learning · Computer Science 2025-03-14 Alceu Bissoto , Trung-Dung Hoang , Tim Flühmann , Susu Sun , Christian F. Baumgartner , Lisa M. Koch

A central obstacle in the objective assessment of treatment effect (TE) estimators in randomized control trials (RCTs) is the lack of ground truth (or validation set) to test their performance. In this paper, we propose a novel…

In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk estimates and its impact on treatment decisions. For overparameterized models, now standard in…

Machine Learning · Computer Science 2026-04-16 Elizabeth W. Miller , Jeffrey D. Blume

Artificial intelligence (AI) is increasingly used to support prognosis in Alzheimer's disease (AD), but adoption remains limited due to a lack of transparency and interpretability, particularly for long-term predictions where uncertainty is…

Human-Computer Interaction · Computer Science 2026-02-03 Jonatan Reyes , Mina Massoumi , Anil Ufuk Batmaz , Marta Kersten-Oertel