English
Related papers

Related papers: Significativity Indices for Agreement Values

200 papers

Studies of the contextual and linguistic factors that constrain discourse phenomena such as reference are coming to depend increasingly on annotated language corpora. In preparing the corpora, it is important to evaluate the reliability of…

cmp-lg · Computer Science 2008-02-03 Rebecca J. Passonneau

In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters…

Machine Learning · Statistics 2019-01-08 Matthijs J. Warrens , Hanneke van der Hoef

Method comparison studies are essential for development in medical and clinical fields. These studies often compare a cheaper, faster, or less invasive measuring method with a widely used one to see if they have sufficient agreement for…

Methodology · Statistics 2019-06-27 Wei Wang , Nan Lin , Jordan D. Oberhaus , Michael S. Avidan

Understanding customer buying patterns is of great interest to the retail industry and has shown to benefit a wide variety of goals ranging from managing stocks to implementing loyalty programs. Association rule mining is a common technique…

Databases · Computer Science 2016-03-16 Martin Kirchgessner , Vincent Leroy , Sihem Amer-Yahia , Shashwat Mishra

A novel approach for comparing quality attributes of different products when there is considerable product-related variability is proposed. In such a case, the whole range of possible realizations must be considered. Looking, for example,…

Methodology · Statistics 2024-08-30 Gerhard Gössler , Vera Hofer , Hans Manner , Walter Goessler

Are existing ways of measuring scientific quality reflecting disadvantages of not being part of giant collaborations? How could possible discrimination be avoided? We propose indices defined for each discipline (subfield) and which count…

Digital Libraries · Computer Science 2016-07-08 A. Tawfik

We commonly use agreement measures to assess the utility of judgements made by human annotators in Natural Language Processing (NLP) tasks. While inter-annotator agreement is frequently used as an indication of label reliability by…

Computation and Language · Computer Science 2025-10-21 Gavin Abercrombie , Tanvi Dinkar , Amanda Cercas Curry , Verena Rieser , Dirk Hovy

Classifier specific (CS) and classifier agnostic (CA) feature importance methods are widely used (often interchangeably) by prior studies to derive feature importance ranks from a defect classifier. However, different feature importance…

Machine Learning · Computer Science 2022-02-08 Gopi Krishnan Rajbahadur , Shaowei Wang , Yasutaka Kamei , Ahmed E. Hassan

As Language Model (LM) capabilities advance, evaluating and supervising them at scale is getting harder for humans. There is hope that other language models can automate both these tasks, which we refer to as ''AI Oversight''. We study how…

We present methods for evaluating human and automatic taggers that extend current practice in three ways. First, we show how to evaluate taggers that assign multiple tags to each test instance, even if they do not assign probabilities.…

Computation and Language · Computer Science 2007-05-23 I. Dan Melamed , Philip Resnik

Currently, computational linguists and cognitive scientists working in the area of discourse and dialogue argue that their subjective judgments are reliable using several different statistics, none of which are easily interpretable or…

cmp-lg · Computer Science 2008-02-03 Jean Carletta

A citation-based indicator for interdisciplinarity has been missing hitherto among the set of available journal indicators. In this study, we investigate network indicators (betweenness centrality), journal indicators (Shannon entropy, the…

Digital Libraries · Computer Science 2010-09-22 Loet Leydesdorff , Ismael Rafols

In an interlaboratory key comparison, a data analysis procedure for this comparison was proposed and recommended by CIPM [1, 2, 3], therein the degrees of equivalence of measurement standards of the laboratories participated in the…

Other Statistics · Statistics 2011-12-14 Thang H. Le , Nguyen D. Do

The large size and complex decision mechanisms of state-of-the-art text classifiers make it difficult for humans to understand their predictions, leading to a potential lack of trust by the users. These issues have led to the adoption of…

Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance,…

Machine Learning · Statistics 2018-10-30 Heinrich Jiang , Been Kim , Melody Y. Guan , Maya Gupta

Large Language Models (LLMs) are increasingly used to annotate learning interactions, yet concerns about reliability limit their utility. We test whether verification-oriented orchestration-prompting models to check their own labels…

Artificial Intelligence · Computer Science 2026-01-29 Bakhtawar Ahtisham , Kirk Vanacore , Jinsook Lee , Zhuqian Zhou , Doug Pietrzak , Rene F. Kizilcec

In evidence-based medicine, relevance of medical literature is determined by predefined relevance conditions. The conditions are defined based on PICO elements, namely, Patient, Intervention, Comparator, and Outcome. Hence, PICO annotations…

Information Retrieval · Computer Science 2023-04-26 Grace E. Lee , Aixin Sun

A popular approach to unveiling the black box of neural NLP models is to leverage saliency methods, which assign scalar importance scores to each input component. A common practice for evaluating whether an interpretability method is…

Computation and Language · Computer Science 2023-05-12 Josip Jukić , Martin Tutek , Jan Šnajder

The integration of artificial intelligence (AI), particularly Convolutional Neural Networks (CNNs), into dermatological diagnosis demonstrates substantial clinical potential. While existing literature predominantly benchmarks algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Loris Cino , Pier Luigi Mazzeo , Alessandro Martella , Giulia Radi , Renato Rossi , Cosimo Distante

The problem of identifying to which of a given set of classes objects belong is ubiquitous, occurring in many research domains and application areas, including medical diagnosis, financial decision making, online commerce, and national…

Machine Learning · Computer Science 2024-09-20 David J. Hand , Peter Christen , Sumayya Ziyad