中文
相关论文

相关论文: Reliability Gaps Between Groups in COMPAS Dataset

200 篇论文

The rapid entry of machine learning approaches in our daily activities and high-stakes domains demands transparency and scrutiny of their fairness and reliability. To help gauge machine learning models' robustness, research typically…

机器学习 · 计算机科学 2023-09-28 Oana Inel , Tim Draws , Lora Aroyo

Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to…

数字图书馆 · 计算机科学 2026-02-06 Lutz Bornmann , Christian Leibel

Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and…

人工智能 · 计算机科学 2025-01-03 Binxia Xu , Antonis Bikakis , Daniel Onah , Andreas Vlachidis , Luke Dickens

Invariance-based randomization tests -- such as permutation tests, rotation tests, or sign changes -- are an important and widely used class of statistical methods. They allow drawing inferences under weak assumptions on the data…

统计理论 · 数学 2022-05-31 Edgar Dobriban

Conformal prediction is a statistically rigorous method for quantifying uncertainty in models by having them output sets of predictions, with larger sets indicating more uncertainty. However, prediction sets are not inherently actionable;…

机器学习 · 计算机科学 2025-02-17 Jesse C. Cresswell , Bhargava Kumar , Yi Sui , Mouloud Belbahri

Raking is widely used in categorical data modeling and survey practice but faced with methodological and computational challenges. We develop a Bayesian paradigm for raking by incorporating the marginal constraints as a prior distribution…

统计方法学 · 统计学 2020-06-24 Yajuan Si , Peigen Zhou

Recently, there has been much research into modelling of uncertainty in human perception through Fuzzy Sets (FSs). Most of this research has focused on allowing respondents to express their (intra) uncertainty using intervals. Here,…

人工智能 · 计算机科学 2019-07-11 Javier Navarro , Christian Wagner

Rating prediction is an important application, and a popular research topic in collaborative filtering. However, both the validity of learning algorithms, and the validity of standard testing procedures rest on the assumption that missing…

机器学习 · 计算机科学 2012-06-26 Benjamin Marlin , Richard S. Zemel , Sam Roweis , Malcolm Slaney

Network analysis provides powerful tools to learn about a variety of social systems. However, most analyses implicitly assume that the considered relational data is error-free, reliable and accurately reflects the system to be analysed.…

社会与信息网络 · 计算机科学 2022-01-12 Felix I. Stamm , Leonie Neuhäuser , Florian Lemmerich , Michael T. Schaub , Markus Strohmaier

Deploying small language models (7-9B parameters) as autonomous agents requires trust in their reasoning, not just their outputs. We reveal a critical reliability crisis: 50-69\% of correct answers from these models contain fundamentally…

机器学习 · 计算机科学 2026-01-05 Laksh Advani

Valid causal inference in observational studies often requires controlling for confounders. However, in practice measurements of confounders may be noisy, and can lead to biased estimates of causal effects. We show that we can reduce the…

机器学习 · 统计学 2018-06-05 Nathan Kallus , Xiaojie Mao , Madeleine Udell

In this paper, we consider a theoretical model for injecting data bias, namely, under-representation and label bias (Blum & Stangl, 2019). We empirically study the effect of varying data biases on the accuracy and fairness of fair…

机器学习 · 计算机科学 2023-12-12 Mohit Sharma , Amit Deshpande , Rajiv Ratn Shah

Machine learning (ML) models used in prediction and classification tasks may display performance disparities across population groups determined by sensitive attributes (e.g., race, sex, age). We consider the problem of evaluating the…

机器学习 · 计算机科学 2024-05-28 Lucas Monteiro Paes , Ananda Theertha Suresh , Alex Beutel , Flavio P. Calmon , Ahmad Beirami

In this paper, we examine the statistical soundness of comparative assessments within the field of recommender systems in terms of reliability and human uncertainty. From a controlled experiment, we get the insight that users provide…

人机交互 · 计算机科学 2017-06-28 Kevin Jasberg , Sergej Sizov

Recommender systems are intrinsically tied to a reliability/coverage dilemma: The more reliable we desire the forecasts, the more conservative the decision will be and thus, the fewer items will be recommended. This causes a detriment to…

信息检索 · 计算机科学 2024-05-22 Diego Pérez-López , Fernando Ortega , Ángel González-Prieto , Jorge Dueñas-Lerín

As machine learning (ML) models gain traction in clinical applications, understanding the impact of clinician and societal biases on ML models is increasingly important. While biases can arise in the labels used for model training, the many…

机器学习 · 计算机科学 2022-08-03 Trenton Chang , Michael W. Sjoding , Jenna Wiens

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

人工智能 · 计算机科学 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

This note is an invited discussion of the article "Confidence Intervals for Nonparametric Empirical Bayes Analysis" by Ignatiadis and Wager. In this discussion, I review some goals of empirical Bayes data analysis and the contribution of…

统计方法学 · 统计学 2022-03-08 Peter Hoff

Biometric recognition is used across a variety of applications from cyber security to border security. Recent research has focused on ensuring biometric performance (false negatives and false positives) is fair across demographic groups.…

统计方法学 · 统计学 2022-08-24 Michael Schuckers , Sandip Purnapatra , Kaniz Fatima , Daqing Hou , Stephanie Schuckers

Machine-learning systems used in survey-based social measurement require uncertainty estimates that are reliable across population subgroups, not merely valid in aggregate. We study ordinal conformal prediction for five-level AI-attitude…

统计方法学 · 统计学 2026-05-08 Amir Rafe , Subasish Das