中文
相关论文

相关论文: Applying Inter-rater Reliability and Agreement in …

200 篇论文

Grounded Theory (GT), a sociological research method designed to study social phenomena, is increasingly being used to investigate the human and social aspects of software engineering (SE). However, being written by and for sociologists, GT…

软件工程 · 计算机科学 2021-09-10 Rashina Hoda

Generative Artificial Intelligence (GenAI) is now widespread in education, yet the efficacy of GenAI systems remains constrained by the quality and interpretation of the labeled data used to train and evaluate them. Studies commonly report…

计算机与社会 · 计算机科学 2026-04-01 Danielle R. Thomas , Conrad Borchers , Kirk P. Vanacore , Kenneth R. Koedinger , René F. Kizilcec

Measurement of the interrater agreement (IRA) is critical in various disciplines. To correct for potential confounding chance agreement in IRA, Cohen's kappa and many other methods have been proposed. However, owing to the varied strategies…

统计方法学 · 统计学 2024-02-14 Zizhong Tian , Vernon M. Chinchilli , Chan Shen , Shouhao Zhou

Humans can be notoriously imperfect evaluators. They are often biased, unreliable, and unfit to define "ground truth." Yet, given the surging need to produce large amounts of training data in educational applications using AI, traditional…

人工智能 · 计算机科学 2025-08-04 Danielle R. Thomas , Conrad Borchers , Kenneth R. Koedinger

In recent years, the research on empirical software engineering that uses qualitative data analysis (e.g., cases studies, interview surveys, and grounded theory studies) is increasing. However, most of this research does not deep into the…

软件工程 · 计算机科学 2025-09-23 Ángel González-Prieto , Jorge Perez , Jessica Diaz , Daniel López-Fernández

Inter-rater reliability (IRR) is one of the commonly used tools for assessing the quality of ratings from multiple raters. However, applicant selection procedures based on ratings from multiple raters usually result in a binary outcome; the…

统计方法学 · 统计学 2025-06-17 František Bartoš , Patrícia Martinková

As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates…

人工智能 · 计算机科学 2026-05-19 Rohit Patel , Alexandre Rezende , Steven McClain

Alpha-based performance evaluation may fail to capture correlated residuals due to model errors. This paper proposes using the Generalized Information Ratio (GIR) to measure performance under misspecified benchmarks. Motivated by the…

投资组合管理 · 定量金融 2018-04-24 Zhongzhi Lawrence He

Since the inception of crowdsourcing, aggregation has been a common strategy for dealing with unreliable data. Aggregate ratings are more reliable than individual ones. However, many natural language processing (NLP) applications that rely…

人工智能 · 计算机科学 2022-03-25 Ka Wong , Praveen Paritosh

Although agreement between annotators has been studied in the past from a statistical viewpoint, little work has attempted to quantify the extent to which this phenomenon affects the evaluation of computer vision (CV) object detection…

计算机视觉与模式识别 · 计算机科学 2018-10-26 Thomas A. Lampert , André Stumpf , Pierre Gançarski

Generating grounded and trustworthy responses remains a key challenge for large language models (LLMs). While retrieval-augmented generation (RAG) with citation-based grounding holds promise, instruction-tuned models frequently fail even in…

计算与语言 · 计算机科学 2025-06-19 Shang Hong Sim , Tej Deep Pala , Vernon Toh , Hai Leong Chieu , Amir Zadeh , Chuan Li , Navonil Majumder , Soujanya Poria

Since the public release of Chat Generative Pre-Trained Transformer (ChatGPT), extensive discourse has emerged concerning the potential advantages and challenges of integrating Generative Artificial Intelligence (GenAI) into education. In…

计算机与社会 · 计算机科学 2024-07-30 Jan-Erik Kalmus , Anastasija Nikiforova

Penetration testing (PT) is an efficient network testing and vulnerability mining tool by simulating a hacker's attack for valuable information applied in some areas. Compared with manual PT, intelligent PT has become a dominating…

密码学与安全 · 计算机科学 2022-04-06 Jinyin Chen , Shulong Hu , Haibin Zheng , Changyou Xing , Guomin Zhang

The availability of big data has significantly influenced the possibilities and methodological choices for conducting large-scale behavioural and social science research. In the context of qualitative data analysis, a major challenge is…

人机交互 · 计算机科学 2025-06-09 Lama Alqazlan , Zheng Fang , Michael Castelle , Rob Procter

Machine learning models are typically evaluated by computing similarity with reference annotations and trained by maximizing similarity with such. Especially in the biomedical domain, annotations are subjective and suffer from low inter-…

How to generate the ground-truth (GT) image is a critical issue for training realistic image super-resolution (Real-ISR) models. Existing methods mostly take a set of high-resolution (HR) images as GTs and apply various degradations to…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Du Chen , Jie Liang , Xindong Zhang , Ming Liu , Hui Zeng , Lei Zhang

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable…

人工智能 · 计算机科学 2026-02-03 Junhyuk Choi , Sohhyung Park , Chanhee Cho , Hyeonchu Park , Bugeun Kim

We introduce an Integrative Ranking and Thresholding (IRT) framework for fusing evidence from multiple testing procedures. The key innovation is a method that transforms binary testing decisions into compound $e-$values, enabling the…

统计方法学 · 统计学 2025-09-04 Trambak Banerjee , Bowen Gang , Jianliang He

Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature…

机器学习 · 计算机科学 2026-05-11 Hyunjun Na , Donghwan Lee

Regression discontinuity (RD) designs with multiple running variables arise in a growing number of empirical applications, including geographic boundaries and multi-score assignment rules. Although recent methodological work has extended…

计量经济学 · 经济学 2026-02-04 Artem Samiahulin
‹ 上一页 1 2 3 10 下一页 ›