中文
相关论文

相关论文: Measuring agreement among several raters classifyi…

200 篇论文

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as ''AI Oversight''. We study how…

In recent years, quadratic weighted kappa has been growing in popularity in the machine learning community as an evaluation metric in domains where the target labels to be predicted are drawn from integer ratings, usually obtained from…

机器学习 · 计算机科学 2015-12-08 David Vaughn , Derek Justice

Multiple raters are often needed to be used interchangeably in practice for measurement or evaluation. Assessing agreement among these multiple raters via agreement indices are necessary before their participation. While the intuitively…

统计方法学 · 统计学 2020-06-09 Tongrong Wang , Huiman X. Barnhart

We study the connection between kappa calculus and probabilistic reasoning in diagnosis applications. Specifically, we abstract a probabilistic belief network for diagnosing faults into a kappa network and compare the ordering of faults…

人工智能 · 计算机科学 2013-02-28 Adnan Darwiche , Moises Goldszmidt

In the $\varepsilon$-Consensus-Halving problem, a fundamental problem in fair division, there are $n$ agents with valuations over the interval $[0,1]$, and the goal is to divide the interval into pieces and assign a label "$+$" or "$-$" to…

计算复杂性 · 计算机科学 2023-04-26 Aris Filos-Ratsikas , Alexandros Hollender , Katerina Sotiraki , Manolis Zampetakis

We propose an evaluation framework for class probability estimates (CPEs) in the presence of label uncertainty, which is commonly observed as diagnosis disagreement between experts in the medical domain. We also formalize evaluation metrics…

机器学习 · 统计学 2021-03-23 Takahiro Mimori , Keiko Sasada , Hirotaka Matsui , Issei Sato

Qualitative analysis of open-ended survey responses is a commonly-used research method in the social sciences, but traditional coding approaches are often time-consuming and prone to inconsistency. Existing solutions from Natural Language…

We formulate three generalized Bayesian models for analyzing interrater and intrarater reliability in the presence of multilevel data. Stan implementations of these models provide new estimates of interrater and intrarater reliability. We…

统计方法学 · 统计学 2024-07-18 Nour Hawila , Arthur Berg

Recently Mart\'in Andr\'es and \'Alvarez Hern\'andez (2024) have proposed new estimators of various kappa coefficients. These estimators are based on the unbiased estimator of the expected index of agreement of each population coefficient.…

统计理论 · 数学 2024-04-10 Antonio Martín Andrés , María Álvarez Hernández

The multi-class prediction had gained popularity over recent years. Thus measuring fit goodness becomes a cardinal question that researchers often have to deal with. Several metrics are commonly used for this task. However, when one has to…

机器学习 · 计算机科学 2022-08-12 Uri Itai , Natan Katz

Causal or unconfounded descriptive comparisons between multiple groups are common in observational studies. Motivated from a racial disparity study in health services research, we propose a unified propensity score weighting framework, the…

统计方法学 · 统计学 2019-07-10 Fan Li , Fan Li

Rating procedure is crucial in many applied fields (e.g., educational, clinical, emergency). It implies that a rater (e.g., teacher, doctor) rates a subject (e.g., student, doctor) on a rating scale. Given raters variability, several…

统计方法学 · 统计学 2026-01-14 Giuseppe Mignemi , Ioanna Manolopoulou

Numerous algorithms have been developed for Conditional Average Treatment Effect (CATE) estimation. In this paper, we first highlight a common issue where many algorithms exhibit inconsistent learning behavior for the same instance across…

机器学习 · 计算机科学 2025-07-08 Yi-Fu Fu , Keng-Te Liao , Shou-De Lin

Multi-label (ML) data deals with multiple classes associated with individual samples at the same time. This leads to the co-occurrence of several classes repeatedly, which indicates some existing correlation among them. In this article, the…

机器学习 · 计算机科学 2021-09-23 Anwesha Law , Ashish Ghosh

This paper appraises the concordance between bibliometrics and peer review, by drawing evidence from the data of two experiments realized by the Italian governmental agency for research evaluation. The experiments were performed for…

应用统计 · 统计学 2020-12-15 Alberto Baccini , Lucio Barabesi , Giuseppe De Nicolao

We present methods for evaluating human and automatic taggers that extend current practice in three ways. First, we show how to evaluate taggers that assign multiple tags to each test instance, even if they do not assign probabilities.…

计算与语言 · 计算机科学 2007-05-23 I. Dan Melamed , Philip Resnik

Classification and differentiation of small pathological objects may greatly vary among human raters due to differences in training, expertise and their consistency over time. In a radiological setting, objects commonly have high…

In directional statistics, the von Mises distribution is a key element in the analysis of circular data. While there is a general agreement regarding the estimation of its location parameter $\mu$, several methods have been proposed to…

应用统计 · 统计学 2024-09-10 Guillaume Marrelec , Alain Giron

Multi-view data, that is matched sets of measurements on the same subjects, have become increasingly common with advances in multi-omics technology. Often, it is of interest to find associations between the views that are related to the…

机器学习 · 统计学 2020-10-02 Yunfeng Zhang , Irina Gaynanova

Binary classification is a task that involves the classification of data into one of two distinct classes. It is widely utilized in various fields. However, conventional classifiers tend to make overconfident predictions for data that…

机器学习 · 计算机科学 2025-03-13 Shoma Yokura , Akihisa Ichiki