中文
相关论文

相关论文: Reliability Gaps Between Groups in COMPAS Dataset

200 篇论文

Recidivism prediction instruments provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the country,…

应用统计 · 统计学 2016-10-25 Alexandra Chouldechova

The abundance of data produced daily from large variety of sources has boosted the need of novel approaches on causal inference analysis from observational data. Observational data often contain noisy or missing entries. Moreover, causal…

统计方法学 · 统计学 2017-03-14 Fani Tsapeli , Peter Tino , Mirco Musolesi

Explainable Artificial Intelligence (XAI) methods (SHAP, LIME) are increasingly adopted to interpret models in high-stakes businesses. However, the credibility of these explanations, their stability under realistic data perturbations,…

人工智能 · 计算机科学 2026-03-06 Alin-Gabriel Vaduva , Simona-Vasilica Oprea , Adela Bara

Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement noise. Current…

机器学习 · 计算机科学 2023-10-10 Michael Hagmann , Philipp Meier , Stefan Riezler

Datagaps are ubiquitous in real world observational data. Quantifying nonlinearity in data having gaps can be challenging. Reported research points out that interpolation can affect nonlinear quantifiers adversely, artificially introducing…

混沌动力学 · 物理学 2018-09-05 Sandip V. George , G. Ambika

Sentiment Analysis Systems (SASs) are data-driven Artificial Intelligence (AI) systems that, given a piece of text, assign one or more numbers conveying the polarity and emotional intensity expressed in the input. Like other automatic…

人工智能 · 计算机科学 2023-02-07 Kausik Lakkaraju , Biplav Srivastava , Marco Valtorta

Differences in data distributions between demographic groups, known as the problem of infra-marginality, complicate how people evaluate fairness in machine learning models. We present a user study with 85 participants in a hypothetical…

人机交互 · 计算机科学 2026-03-09 Schrasing Tong , Minseok Jung , Ilaria Liccardi , Lalana Kagal

For multiple reasons -- such as avoiding overtraining from one data set or because of having received numerical estimates for some parameters in a model from an alternative source -- it is sometimes useful to divide a model's parameters…

统计方法学 · 统计学 2024-06-26 Yunrong Wan

Instrumental variable methods are among the most commonly used causal inference approaches to deal with unmeasured confounders in observational studies. The presence of invalid instruments is the primary concern for practical applications,…

统计方法学 · 统计学 2023-04-18 Zijian Guo

How reliably can we trust the scores obtained from social bias benchmarks as faithful indicators of problematic social biases in a given language model? In this work, we study this question by contrasting social biases with non-social…

计算与语言 · 计算机科学 2023-06-21 Nikil Roashan Selvam , Sunipa Dev , Daniel Khashabi , Tushar Khot , Kai-Wei Chang

Systematic error, which is not determined by chance, often refers to the inaccuracy (involving either the observation or measurement process) inherent to a system. In this paper, we exhibit some long-neglected but frequent-happening…

计算机视觉与模式识别 · 计算机科学 2021-09-03 Yan Wang , Yuhang Li , Ruihao Gong

This paper addresses the issue of {\sc pos} tagger evaluation. Such evaluation is usually performed by comparing the tagger output with a reference test corpus, which is assumed to be error-free. Currently used corpora contain noise which…

计算与语言 · 计算机科学 2007-05-23 L. Padro , L. Marquez

Errors in labels obtained via human annotation adversely affect a model's performance. Existing approaches propose ways to mitigate the effect of label error on a model's downstream accuracy, yet little is known about its impact on a…

机器学习 · 计算机科学 2023-10-05 Julius Adebayo , Melissa Hall , Bowen Yu , Bobbie Chern

Inter-rater reliability (IRR) is one of the commonly used tools for assessing the quality of ratings from multiple raters. However, applicant selection procedures based on ratings from multiple raters usually result in a binary outcome; the…

统计方法学 · 统计学 2025-06-17 František Bartoš , Patrícia Martinková

When using machine learning to aid decision-making, it is critical to ensure that an algorithmic decision is fair and does not discriminate against specific individuals/groups, particularly those from underprivileged populations. Existing…

机器学习 · 计算机科学 2024-11-20 Yifei Wang , Zhengyang Zhou , Liqin Wang , John Laurentiev , Peter Hou , Li Zhou , Pengyu Hong

Risk Assessment Instruments (RAIs) are widely used to forecast adverse outcomes in domains such as healthcare and criminal justice. RAIs are commonly trained on observational data and are optimized to predict observable outcomes rather than…

统计方法学 · 统计学 2021-04-06 Alan Mishler , Niccolò Dalmasso

Recommender systems rely heavily on user feedback to learn effective user and item representations. Despite their widespread adoption, limited attention has been given to the uncertainty inherent in the feedback used to train these systems.…

信息检索 · 计算机科学 2025-05-06 Bruno Sguerra , Viet-Anh Tran , Romain Hennequin , Manuel Moussallam

This paper demonstrates a methodology for examining the accuracy of uncertain inference systems (UIS), after their parameters have been optimized, and does so for several common UIS's. This methodology may be used to test the accuracy when…

人工智能 · 计算机科学 2013-04-11 Ben P. Wise

The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, healthcare decision-making, and automated risk assessment. Existing approaches typically rely…

机器学习 · 计算机科学 2026-05-21 Khalid Adnan Alsayed

Experimental datasets are growing rapidly in size, scope, and detail, but the value of these datasets is limited by unwanted measurement noise. It is therefore tempting to apply analysis techniques that attempt to reduce noise and enhance…

应用统计 · 统计学 2022-07-12 Kendrick Kay