English
Related papers

Related papers: Liberal-Conservative Hierarchies of Intercoder Rel…

200 papers

Inter-coder agreement measures, like Cohen's kappa, correct the relative frequency of agreement between coders to account for agreement which simply occurs by chance. However, in some situations these measures exhibit behavior which make…

Applications · Statistics 2012-08-07 Dirk Schuster

LLMs enable qualitative coding at large scale, but assessing reliability remains challenging where human experts seldom agree. We investigate confidence-diversity calibration as a quality assessment framework for accessible coding tasks…

Machine Learning · Computer Science 2025-08-19 Zhilong Zhao , Yindi Liu

Comparing alternatives in pairs is a very well known technique of ranking creation. The answer to how reliable and trustworthy ranking is depends on the inconsistency of the data from which it was created. There are many indices used for…

Discrete Mathematics · Computer Science 2020-01-28 Konrad Kułakowski , Dawid Talaga

Agreement measures, such as Cohen's kappa or intraclass correlation, gauge the matching between two or more classifiers. They are used in a wide range of contexts from medicine, where they evaluate the effectiveness of medical treatments…

Machine Learning · Computer Science 2025-09-23 Alberto Casagrande , Francesco Fabris , Rossano Girometti , Roberto Pagliarini

Background. The reliability paradox describes the empirical observation that cognitive tasks producing robust group-level effects often yield poor between-individual reliability. Existing approaches rely predominantly on the intraclass…

Methodology · Statistics 2026-05-26 Maria Westrin

Cohen's kappa is a useful measure for agreement between the judges, inter-rater reliability, and also goodness of fit in classification problems. For binary nominal and ordinal data, kappa and correlation are equally applicable. We have…

Methodology · Statistics 2024-04-23 Soumya Sahu , Hakan Demirtas

Deep neural networks, while powerful for image classification, often operate as "black boxes," complicating the understanding of their decision-making processes. Various explanation methods, particularly those generating saliency maps, aim…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Tristan Gomez , Harold Mouchère

To measure the degree of agreement between two observers that independently classify $n$ subjects within $K$ categories, it is common to use different kappa type coefficients, the most common of which is the $\kappa_C$ coefficient (Cohen's…

Statistics Theory · Mathematics 2026-02-24 A. Martín Andrés , M. Álvarez Hernández

Reliable uncertainty quantification is crucial for the trustworthiness of machine learning applications. Inductive Conformal Prediction (ICP) offers a distribution-free framework for generating prediction sets or intervals with…

Machine Learning · Computer Science 2025-06-25 A. A. Balinsky , A. D. Balinsky

Index coding and coded caching are two active research topics in information theory with strong ties to each other. Motivated by the multi-access coded caching problem, we study a new class of structured index coding problems (ICPs) which…

Information Theory · Computer Science 2021-11-17 Kota Srinivas Reddy , Nikhil Karamchandani

Large language models (LLMs) are increasingly deployed for tabular question answering, yet calibration on structured data is largely unstudied. This paper presents the first systematic comparison of five confidence estimation methods across…

Computation and Language · Computer Science 2026-04-15 Lukas Voss

We study higher analogues of the classical independence number on $\omega$. For $\kappa$ regular uncountable, we denote by $i(\kappa)$ the minimal size of a maximal $\kappa$-independent family. We establish ZFC relations between $i(\kappa)$…

Logic · Mathematics 2022-06-10 Vera Fischer , Diana Carolina Montoya

We propose an information-theoretic alternative to the popular Cronbach alpha coefficient of reliability. Particularly suitable for contexts in which instruments are scored on a strictly nonnumeric scale, our proposed index is based on…

Statistics Theory · Mathematics 2015-01-19 Ernest Fokoue , Necla Gunduz

Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite known limitations in their reliability and robustness. Yet how they shape researchers'…

Computation and Language · Computer Science 2026-05-29 Pouya Sadeghi , Anamaria Crisan , Jimmy Lin

Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However, the influence of endorsement source credibility remains underexplored. We investigate…

Computation and Language · Computer Science 2026-05-28 Priyanka Mary Mammen , Emil Joswin , Shankar Venkitachalam

We consider the weakly supervised binary classification problem where the labels are randomly flipped with probability $1- {\alpha}$. Although there exist numerous algorithms for this problem, it remains theoretically unexplored how the…

Machine Learning · Computer Science 2019-07-16 Xinyang Yi , Zhaoran Wang , Zhuoran Yang , Constantine Caramanis , Han Liu

Hyper-partisan misinformation has become a major public concern. In order to examine what type of misinformation label can mitigate hyper-partisan misinformation sharing on social media, we conducted a 4 (label type: algorithm, community,…

Human-Computer Interaction · Computer Science 2023-01-24 Chenyan Jia , Alexander Boltz , Angie Zhang , Anqing Chen , Min Kyung Lee

We study linear preconditioning in Markov chain Monte Carlo. We consider the class of well-conditioned distributions, for which several mixing time bounds depend on the condition number $\kappa$. First we show that well-conditioned…

Computation · Statistics 2024-12-05 Max Hird , Samuel Livingstone

We assessed several agreement coefficients applied in 2x2 contingency tables, which are commonly applied in research due to dicotomization by the conditions of the subjects (e.g., male or female) or by conveniency of the classification…

Methodology · Statistics 2022-04-14 Paulo Sergio Panse Silveira , Jose Oliveira Siqueira

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

Methodology · Statistics 2025-09-22 Filip Moons , Ellen Vandervieren
‹ Prev 1 2 3 10 Next ›