中文
相关论文

相关论文: Better Than Their Reputation? On the Reliability o…

200 篇论文

This paper is about an information retrieval evaluation on three different retrieval-supporting services. All three services were designed to compensate typical problems that arise in metadata-driven Digital Libraries, which are not…

信息检索 · 计算机科学 2010-10-12 Philipp Schaer , Philipp Mayr , Peter Mutschke

Relevance and fairness are two major objectives of recommender systems (RSs). Recent work proposes measures of RS fairness that are either independent from relevance (fairness-only) or conditioned on relevance (joint measures). While…

信息检索 · 计算机科学 2024-05-29 Theresia Veronika Rampisela , Tuukka Ruotsalo , Maria Maistro , Christina Lioma

In recent years, the research on empirical software engineering that uses qualitative data analysis (e.g., cases studies, interview surveys, and grounded theory studies) is increasing. However, most of this research does not deep into the…

软件工程 · 计算机科学 2025-09-23 Ángel González-Prieto , Jorge Perez , Jessica Diaz , Daniel López-Fernández

Recent discussions on alternative facts, fake news, and post truth politics have motivated research on creating technologies that allow people not only to access information, but also to assess the credibility of the information presented…

信息检索 · 计算机科学 2017-08-25 Christina Lioma , Jakob Grue Simonsen , Birger Larsen

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

统计方法学 · 统计学 2025-09-22 Filip Moons , Ellen Vandervieren

LLM-as-a-Judge has emerged as a popular evaluation strategy, where advanced large language models assess generation results in alignment with human instructions. While these models serve as a promising alternative to human annotators, their…

计算与语言 · 计算机科学 2025-05-20 Xiyan Fu , Wei Liu

Relevance judgment of human assessors is inherently subjective and dynamic when evaluation datasets are created for Information Retrieval (IR) systems. However, a small group of experts' relevance judgment results are usually taken as…

信息检索 · 计算机科学 2022-08-09 Dengya Zhu , Shastri L Nimmagadda , Kok Wai Wong , Torsten Reiners

The evaluation of Information Retrieval (IR) systems typically uses query-document pairs with corresponding human-labelled relevance assessments (qrels). These qrels are used to determine if one system is better than another based on…

信息检索 · 计算机科学 2025-07-11 Jack McKechnie , Graham McDonald , Craig Macdonald

The performance of machine learning classification algorithms are evaluated by estimating metrics, often from the confusion matrix, using training data and cross-validation. However, these do not prove that the best possible performance has…

机器学习 · 统计学 2024-03-05 L. Crow , S. J. Watts

Over the last decade there has been increasing concern about the biases embodied in traditional evaluation methods for Natural Language Processing/Learning, particularly methods borrowed from Information Retrieval. Without knowledge of the…

人工智能 · 计算机科学 2015-04-06 David M. W. Powers

The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards…

Qualitative analysis is typically limited to small datasets because it is time-intensive. Moreover, a second human rater is required to ensure reliable findings. Artificial intelligence tools may replace human raters if we demonstrate high…

物理教育 · 物理学 2025-09-03 Nikhil Sanjay Borse , Ravishankar Chatta Subramaniam , N. Sanjay Rebello

Reference texts such as encyclopedias and news articles can manifest biased language when objective reporting is substituted by subjective writing. Existing methods to detect bias mostly rely on annotated data to train machine learning…

计算与语言 · 计算机科学 2021-12-20 Timo Spinde , David Krieger , Manuel Plank , Bela Gipp

Besides position bias, which has been well-studied, trust bias is another type of bias prevalent in user interactions with rankings: users are more likely to click incorrectly w.r.t. their preferences on highly ranked items because they…

信息检索 · 计算机科学 2020-09-10 Ali Vardasbi , Harrie Oosterhuis , Maarten de Rijke

The need to measure the degree of agreement among R raters who independently classify n subjects within K nominal categories is frequent in many scientific areas. The most popular measures are Cohen's kappa (R = 2), Fleiss' kappa, Conger's…

应用统计 · 统计学 2022-02-01 A. Martín Andrés , M. Álvarez Hernández

Large language models (LLMs) are increasingly used to assign document relevance labels in information retrieval pipelines, especially in domains lacking human-labeled data. However, different models often disagree on borderline cases,…

信息检索 · 计算机科学 2025-07-04 William A. Ingram , Bipasha Banerjee , Edward A. Fox

In realistic retrieval settings with large and evolving knowledge bases, the total number of documents relevant to a query is typically unknown, and recall cannot be computed. In this paper, we evaluate several established strategies for…

计算与语言 · 计算机科学 2026-05-08 Shelly Schwartz , Oleg Vasilyev , Randy Sawaya

This study examines a basic assumption of peer review, namely, the idea that there is a consensus on evaluation criteria among peers, which is a necessary condition for the reliability of peer judgements. Empirical evidence indicating that…

社会与信息网络 · 计算机科学 2022-01-10 Sven E. Hug , Michael Ochsner

In any ranking system, the retrieval model outputs a single score for a document based on its belief on how relevant it is to a given search query. While retrieval models have continued to improve with the introduction of increasingly…

信息检索 · 计算机科学 2021-05-12 Daniel Cohen , Bhaskar Mitra , Oleg Lesota , Navid Rekabsaz , Carsten Eickhoff

Inter-Rater quantifies the reliability between multiple raters who evaluate a group of subjects. It calculates the group quantity, Fleiss kappa, and it improves on existing software by keeping information about each user and quantifying how…

其他统计学 · 统计学 2018-09-18 Daniel J. Arenas
‹ 上一页 1 2 3 10 下一页 ›