中文
相关论文

相关论文: Evaluating AI systems under uncertain ground truth…

200 篇论文

Machine learning risks reinforcing biases present in data and, as we argue in this work, in what is absent from data. In healthcare, societal and decision biases shape patterns in missing data, yet the algorithmic fairness implications of…

人工智能 · 计算机科学 2025-03-19 Vincent Jeanselme , Maria De-Arteaga , Zhe Zhang , Jessica Barrett , Brian Tom

The ability to acknowledge the inevitable uncertainty in their knowledge and reasoning is a prerequisite for AI systems to be truly truthful and reliable. In this paper, we present a taxonomy of uncertainty specific to vision-language AI…

人工智能 · 计算机科学 2024-07-03 Khyathi Raghavi Chandu , Linjie Li , Anas Awadalla , Ximing Lu , Jae Sung Park , Jack Hessel , Lijuan Wang , Yejin Choi

We evaluate artificial intelligence (AI) systems without ground truth by exploiting a link between strategic gaming and information loss. Building on established information theory, we analyze which mechanisms resist adversarial…

机器学习 · 计算机科学 2026-05-01 Zachary Robertson , Sanmi Koyejo

Reliable evaluation of AI systems remains a fundamental challenge when ground truth labels are unavailable, particularly for systems generating natural language outputs like AI chat and agent systems. Many of these AI agents and systems…

机器学习 · 统计学 2025-11-05 Kaihua Ding

The deployment of AI systems for welfare benefit allocation allows for accelerated decision-making and faster provision of critical help, but has already led to an increase in unfair benefit denials and false fraud accusations. Collecting…

计算机与社会 · 计算机科学 2024-07-18 Mengchen Dong , Jean-François Bonnefon , Iyad Rahwan

Early detection and rapid intervention of lung cancer are crucial. Nonetheless, ensuring an accurate diagnosis is challenging, as physicians' ability to interpret chest X-rays varies significantly depending on their experience and degree of…

图像与视频处理 · 电气工程与系统科学 2025-08-20 Hyeonjin Choi , Jinse Kim , Dong-yeon Yoo , Ju-sung Sun , Jung-won Lee

This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct…

计算与语言 · 计算机科学 2026-05-27 Shang Wu , Randol Yao

The assessment of process mining techniques using real-life data is often compromised by the lack of ground truth knowledge, the presence of non-essential outliers in system behavior and recording errors in event logs. Using synthetically…

数据库 · 计算机科学 2025-01-27 Dominique Sommers , Natalia Sidorova , Boudewijn van Dongen

In safety-critical applications like medical diagnosis, certainty associated with a model's prediction is just as important as its accuracy. Consequently, uncertainty estimation and reduction play a crucial role. Uncertainty in predictions…

图像与视频处理 · 电气工程与系统科学 2023-09-12 Abhishek Singh Sambyal , Narayanan C. Krishnan , Deepti R. Bathula

Although machine learning (ML) models of AI achieve high performances in medicine, they are not free of errors. Empowering clinicians to identify incorrect model recommendations is crucial for engendering trust in medical AI. Explainable AI…

人工智能 · 计算机科学 2022-12-20 Isil Guzey , Ozlem Ucar , Nukhet Aladag Ciftdemir , Betul Acunas

Medicine and deep learning-based artificial intelligence (AI) engineering represent two distinct fields each with decades of published history. With such history comes a set of terminology that has a specific way in which it is applied.…

Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system identifies, how it prioritizes them, or whether those priorities align with the review…

人工智能 · 计算机科学 2026-04-23 Ming Jin

Artificial intelligence (AI) has demonstrated the ability to extract insights from data, but the issue of fairness remains a concern in high-stakes fields such as healthcare. Despite extensive discussion and efforts in algorithm…

Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels).…

机器学习 · 计算机科学 2021-10-29 Maria Ulan , Welf Löwe , Morgan Ericsson , Anna Wingkvist

The demand for Explainable AI (XAI) has triggered an explosion of methods, producing a landscape so fragmented that we now rely on surveys of surveys. Yet, fundamental challenges persist: conflicting metrics, failed sanity checks, and…

机器学习 · 计算机科学 2026-03-31 Amir-Hossein Karimi

Deep learning algorithms have demonstrated remarkable efficacy in various medical image analysis (MedIA) applications. However, recent research highlights a performance disparity in these algorithms when applied to specific subgroups, such…

计算机视觉与模式识别 · 计算机科学 2024-07-10 Zikang Xu , Jun Li , Qingsong Yao , Han Li , Mingyue Zhao , S. Kevin Zhou

Machine learning (ML) models may suffer from significant performance disparities between patient groups. Identifying such disparities by monitoring performance at a granular level is crucial for safely deploying ML to each patient.…

机器学习 · 计算机科学 2025-03-14 Alceu Bissoto , Trung-Dung Hoang , Tim Flühmann , Susu Sun , Christian F. Baumgartner , Lisa M. Koch

A central obstacle in the objective assessment of treatment effect (TE) estimators in randomized control trials (RCTs) is the lack of ground truth (or validation set) to test their performance. In this paper, we propose a novel…

In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk estimates and its impact on treatment decisions. For overparameterized models, now standard in…

机器学习 · 计算机科学 2026-04-16 Elizabeth W. Miller , Jeffrey D. Blume

Artificial intelligence (AI) is increasingly used to support prognosis in Alzheimer's disease (AD), but adoption remains limited due to a lack of transparency and interpretability, particularly for long-term predictions where uncertainty is…

人机交互 · 计算机科学 2026-02-03 Jonatan Reyes , Mina Massoumi , Anil Ufuk Batmaz , Marta Kersten-Oertel