English
Related papers

Related papers: Significativity Indices for Agreement Values

200 papers

Inter-coder agreement measures, like Cohen's kappa, correct the relative frequency of agreement between coders to account for agreement which simply occurs by chance. However, in some situations these measures exhibit behavior which make…

Applications · Statistics 2012-08-07 Dirk Schuster

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

Methodology · Statistics 2025-09-22 Filip Moons , Ellen Vandervieren

To measure the degree of agreement between two observers that independently classify $n$ subjects within $K$ categories, it is common to use different kappa type coefficients, the most common of which is the $\kappa_C$ coefficient (Cohen's…

Statistics Theory · Mathematics 2026-02-24 A. Martín Andrés , M. Álvarez Hernández

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the…

Methodology · Statistics 2026-01-16 William L. Lippitt , Edward J. Bedrick , Nichole E. Carlson

Agreement measures are useful to both compare different evaluations of the same diagnostic outcomes and validate new rating systems or devices. Information Agreement (IA) is an information-theoretic-based agreement measure introduced to…

Information Theory · Computer Science 2020-08-27 Alberto Casagrande , Francesco Fabris , Rossano Girometti

The need to measure the degree of agreement among R raters who independently classify n subjects within K nominal categories is frequent in many scientific areas. The most popular measures are Cohen's kappa (R = 2), Fleiss' kappa, Conger's…

Applications · Statistics 2022-02-01 A. Martín Andrés , M. Álvarez Hernández

Measurement of the interrater agreement (IRA) is critical in various disciplines. To correct for potential confounding chance agreement in IRA, Cohen's kappa and many other methods have been proposed. However, owing to the varied strategies…

Methodology · Statistics 2024-02-14 Zizhong Tian , Vernon M. Chinchilli , Chan Shen , Shouhao Zhou

We assessed several agreement coefficients applied in 2x2 contingency tables, which are commonly applied in research due to dicotomization by the conditions of the subjects (e.g., male or female) or by conveniency of the classification…

Methodology · Statistics 2022-04-14 Paulo Sergio Panse Silveira , Jose Oliveira Siqueira

Cohen's kappa is a useful measure for agreement between the judges, inter-rater reliability, and also goodness of fit in classification problems. For binary nominal and ordinal data, kappa and correlation are equally applicable. We have…

Methodology · Statistics 2024-04-23 Soumya Sahu , Hakan Demirtas

Classification is a machine learning method used in many practical applications: text mining, handwritten character recognition, face recognition, pattern classification, scene labeling, computer vision, natural langage processing. A…

Machine Learning · Computer Science 2025-11-05 Doulaye Dembélé

The comparison of alternative rankings of a set of items is a general and prominent task in applied statistics. Predictor variables are ranked according to magnitude of association with an outcome, prediction models rank subjects according…

The weighted kappa coefficient of a binary diagnostic test is a measure of the beyond-chance agreement between the diagnostic test and the gold standard, and depends on the sensitivity and specificity of the diagnostic test, on the disease…

Other Statistics · Statistics 2024-09-02 Jose Antonio Roldan-Nofuentes , Saad bouh Sidaty-regad

When annotators label data, a key metric for quality assurance is inter-annotator agreement (IAA): the extent to which annotators agree on their labels. Though many IAA measures exist for simple categorical and ordinal labeling tasks,…

Computation and Language · Computer Science 2022-12-20 Alexander Braylan , Omar Alonso , Matthew Lease

While advanced classifiers have been increasingly used in real-world safety-critical applications, how to properly evaluate the black-box models given specific human values remains a concern in the community. Such human values include…

Machine Learning · Computer Science 2024-03-14 Yanyun Wang , Dehui Du , Yuanhao Liu

Agreement coefficients provide a fundamental framework for quantifying the concordance between two or more measurement methods applied to the same continuous variable. Unlike correlation, which measures the strength of a linear…

Methodology · Statistics 2026-04-28 Ronny Vallejos

Human annotation remains the foundation of reliable and interpretable data in Natural Language Processing (NLP). As annotation and evaluation tasks continue to expand, from categorical labelling to segmentation, subjective judgment, and…

Computation and Language · Computer Science 2026-04-02 Joseph James

Qualitative research faces a critical reliability challenge: traditional inter-rater agreement methods require multiple human coders, are time-intensive, and often yield moderate consistency. We present a multi-perspective validation…

Computation and Language · Computer Science 2026-02-17 Nilesh Jain , Hyungil Suh , Seyi Adeyinka , Leor Roseman , Aza Allsop

Complex assignments typically consist of open-ended questions with large and diverse content in the context of both classroom and online graduate programs. With the sheer scale of these programs comes a variety of problems in peer and…

Computation and Language · Computer Science 2020-03-17 Manikandan Ravikiran

Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects unclear criteria, collapsed distinctions, or legitimate plurality. We propose a…

Computation and Language · Computer Science 2026-05-01 Nisrine Rair , Alban Goupil , Valeriu Vrabie , Emmanuel Chochoy

During the Italian research assessment exercise, the national agency ANVUR performed an experiment to assess agreement between grades attributed to journal articles by informed peer review (IR) and by bibliometrics. A sample of articles was…

Digital Libraries · Computer Science 2016-03-25 Alberto Baccini , Giuseppe De Nicolao
‹ Prev 1 2 3 10 Next ›