English
Related papers

Related papers: Cross-replication Reliability -- An Empirical Appr…

200 papers

Measurement of the interrater agreement (IRA) is critical in various disciplines. To correct for potential confounding chance agreement in IRA, Cohen's kappa and many other methods have been proposed. However, owing to the varied strategies…

Methodology · Statistics 2024-02-14 Zizhong Tian , Vernon M. Chinchilli , Chan Shen , Shouhao Zhou

This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology. QRA produces a single score estimating the degree of reproducibility of a given…

Computation and Language · Computer Science 2022-04-13 Anya Belz , Maja Popović , Simon Mille

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) applications that rely…

Artificial Intelligence · Computer Science 2022-03-25 Ka Wong , Praveen Paritosh

Inter-rater reliability (IRR) is one of the commonly used tools for assessing the quality of ratings from multiple raters. However, applicant selection procedures based on ratings from multiple raters usually result in a binary outcome; the…

Methodology · Statistics 2025-06-17 František Bartoš , Patrícia Martinková

Face swapping has become a prominent research area in computer vision and image processing due to rapid technological advancements. The metric of measuring the quality in most face swapping methods relies on several distances between the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Xinghui Zhou , Wenbo Zhou , Tianyi Wei , Shen Chen , Taiping Yao , Shouhong Ding , Weiming Zhang , Nenghai Yu

Inter-coder agreement measures, like Cohen's kappa, correct the relative frequency of agreement between coders to account for agreement which simply occurs by chance. However, in some situations these measures exhibit behavior which make…

Applications · Statistics 2012-08-07 Dirk Schuster

Inter-Rater quantifies the reliability between multiple raters who evaluate a group of subjects. It calculates the group quantity, Fleiss kappa, and it improves on existing software by keeping information about each user and quantifying how…

Other Statistics · Statistics 2018-09-18 Daniel J. Arenas

As super-resolution (SR) techniques advance, we observe a growing distrust of evaluation metrics in recent SR research. An inconsistency often emerges between certain evaluation criteria and human perceptual preference. Although current SR…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Shaolin Su , Josep M. Rocafort , Danna Xue , David Serrano-Lozano , Lei Sun , Javier Vazquez-Corral

Inter-rater reliability (IRR), which is a prerequisite of high-quality ratings and assessments, may be affected by contextual variables such as the rater's or ratee's gender, major, or experience. Identification of such heterogeneity…

Methodology · Statistics 2023-02-17 Patrícia Martinková , František Bartoš , Marek Brabec

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Ding Jiang , Mang Ye

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

Methodology · Statistics 2025-09-22 Filip Moons , Ellen Vandervieren

Reject Inference (RI) methods aim to address sample bias by inferring missing repayment data for rejected credit applicants. Traditional approaches often assume that the behavior of rejected clients can be extrapolated from accepted…

Machine Learning · Computer Science 2025-10-16 Athyrson Machado Ribeiro , Marcos Medeiros Raimundo

We formulate three generalized Bayesian models for analyzing interrater and intrarater reliability in the presence of multilevel data. Stan implementations of these models provide new estimates of interrater and intrarater reliability. We…

Methodology · Statistics 2024-07-18 Nour Hawila , Arthur Berg

Reproduction studies reported in NLP provide individual data points which in combination indicate worryingly low levels of reproducibility in the field. Because each reproduction study reports quantitative conclusions based on its own,…

Computation and Language · Computer Science 2025-05-26 Anya Belz

Reproducibility is essential to reliable scientific discovery in high-throughput experiments. In this work we propose a unified approach to measure the reproducibility of findings identified from replicate experiments and identify putative…

Applications · Statistics 2011-10-24 Qunhua Li , James B. Brown , Haiyan Huang , Peter J. Bickel

In recent years, the qualitative research on empirical software engineering that applies Grounded Theory is increasing. Grounded Theory (GT) is a technique for developing theory inductively e iteratively from qualitative data based on…

Software Engineering · Computer Science 2021-07-27 Jessica Díaz , Jorge Pérez , Carolina Gallardo , Ángel González-Prieto

Studies of the contextual and linguistic factors that constrain discourse phenomena such as reference are coming to depend increasingly on annotated language corpora. In preparing the corpora, it is important to evaluate the reliability of…

cmp-lg · Computer Science 2008-02-03 Rebecca J. Passonneau

Scale-invariance is an open problem in many computer vision subfields. For example, object labels should remain constant across scales, yet model predictions diverge in many cases. This problem gets harder for tasks where the ground-truth…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Oliver Wiedemann , Vlad Hosu , Shaolin Su , Dietmar Saupe

With the aim of matching a pair of instances from two different modalities, cross modality mapping has attracted growing attention in the computer vision community. Existing methods usually formulate the mapping function as the similarity…

Computer Vision and Pattern Recognition · Computer Science 2021-03-01 Zun Li , Congyan Lang , Liqian Liang , Tao Wang , Songhe Feng , Jun Wu , Yidong Li

Inspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Rizhao Cai , Haoliang Li , Shiqi Wang , Changsheng Chen , Alex Chichung Kot
‹ Prev 1 2 3 10 Next ›