中文
相关论文

相关论文: Cross-replication Reliability -- An Empirical Appr…

200 篇论文

Measurement of the interrater agreement (IRA) is critical in various disciplines. To correct for potential confounding chance agreement in IRA, Cohen's kappa and many other methods have been proposed. However, owing to the varied strategies…

统计方法学 · 统计学 2024-02-14 Zizhong Tian , Vernon M. Chinchilli , Chan Shen , Shouhao Zhou

This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology. QRA produces a single score estimating the degree of reproducibility of a given…

计算与语言 · 计算机科学 2022-04-13 Anya Belz , Maja Popović , Simon Mille

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) applications that rely…

人工智能 · 计算机科学 2022-03-25 Ka Wong , Praveen Paritosh

Inter-rater reliability (IRR) is one of the commonly used tools for assessing the quality of ratings from multiple raters. However, applicant selection procedures based on ratings from multiple raters usually result in a binary outcome; the…

统计方法学 · 统计学 2025-06-17 František Bartoš , Patrícia Martinková

Face swapping has become a prominent research area in computer vision and image processing due to rapid technological advancements. The metric of measuring the quality in most face swapping methods relies on several distances between the…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Xinghui Zhou , Wenbo Zhou , Tianyi Wei , Shen Chen , Taiping Yao , Shouhong Ding , Weiming Zhang , Nenghai Yu

Inter-coder agreement measures, like Cohen's kappa, correct the relative frequency of agreement between coders to account for agreement which simply occurs by chance. However, in some situations these measures exhibit behavior which make…

应用统计 · 统计学 2012-08-07 Dirk Schuster

Inter-Rater quantifies the reliability between multiple raters who evaluate a group of subjects. It calculates the group quantity, Fleiss kappa, and it improves on existing software by keeping information about each user and quantifying how…

其他统计学 · 统计学 2018-09-18 Daniel J. Arenas

As super-resolution (SR) techniques advance, we observe a growing distrust of evaluation metrics in recent SR research. An inconsistency often emerges between certain evaluation criteria and human perceptual preference. Although current SR…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Shaolin Su , Josep M. Rocafort , Danna Xue , David Serrano-Lozano , Lei Sun , Javier Vazquez-Corral

Inter-rater reliability (IRR), which is a prerequisite of high-quality ratings and assessments, may be affected by contextual variables such as the rater's or ratee's gender, major, or experience. Identification of such heterogeneity…

统计方法学 · 统计学 2023-02-17 Patrícia Martinková , František Bartoš , Marek Brabec

Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have…

计算机视觉与模式识别 · 计算机科学 2023-03-23 Ding Jiang , Mang Ye

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

统计方法学 · 统计学 2025-09-22 Filip Moons , Ellen Vandervieren

Reject Inference (RI) methods aim to address sample bias by inferring missing repayment data for rejected credit applicants. Traditional approaches often assume that the behavior of rejected clients can be extrapolated from accepted…

机器学习 · 计算机科学 2025-10-16 Athyrson Machado Ribeiro , Marcos Medeiros Raimundo

We formulate three generalized Bayesian models for analyzing interrater and intrarater reliability in the presence of multilevel data. Stan implementations of these models provide new estimates of interrater and intrarater reliability. We…

统计方法学 · 统计学 2024-07-18 Nour Hawila , Arthur Berg

Reproduction studies reported in NLP provide individual data points which in combination indicate worryingly low levels of reproducibility in the field. Because each reproduction study reports quantitative conclusions based on its own,…

计算与语言 · 计算机科学 2025-05-26 Anya Belz

Reproducibility is essential to reliable scientific discovery in high-throughput experiments. In this work we propose a unified approach to measure the reproducibility of findings identified from replicate experiments and identify putative…

应用统计 · 统计学 2011-10-24 Qunhua Li , James B. Brown , Haiyan Huang , Peter J. Bickel

In recent years, the qualitative research on empirical software engineering that applies Grounded Theory is increasing. Grounded Theory (GT) is a technique for developing theory inductively e iteratively from qualitative data based on…

软件工程 · 计算机科学 2021-07-27 Jessica Díaz , Jorge Pérez , Carolina Gallardo , Ángel González-Prieto

Studies of the contextual and linguistic factors that constrain discourse phenomena such as reference are coming to depend increasingly on annotated language corpora. In preparing the corpora, it is important to evaluate the reliability of…

cmp-lg · 计算机科学 2008-02-03 Rebecca J. Passonneau

Scale-invariance is an open problem in many computer vision subfields. For example, object labels should remain constant across scales, yet model predictions diverge in many cases. This problem gets harder for tasks where the ground-truth…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Oliver Wiedemann , Vlad Hosu , Shaolin Su , Dietmar Saupe

With the aim of matching a pair of instances from two different modalities, cross modality mapping has attracted growing attention in the computer vision community. Existing methods usually formulate the mapping function as the similarity…

计算机视觉与模式识别 · 计算机科学 2021-03-01 Zun Li , Congyan Lang , Liqian Liang , Tao Wang , Songhe Feng , Jun Wu , Yidong Li

Inspired by the philosophy employed by human beings to determine whether a presented face example is genuine or not, i.e., to glance at the example globally first and then carefully observe the local regions to gain more discriminative…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Rizhao Cai , Haoliang Li , Shiqi Wang , Changsheng Chen , Alex Chichung Kot
‹ 上一页 1 2 3 10 下一页 ›