中文
相关论文

相关论文: Assessing agreement on classification tasks: the k…

200 篇论文

We study the connection between kappa calculus and probabilistic reasoning in diagnosis applications. Specifically, we abstract a probabilistic belief network for diagnosing faults into a kappa network and compare the ordering of faults…

人工智能 · 计算机科学 2013-02-28 Adnan Darwiche , Moises Goldszmidt

To measure the degree of agreement between two observers that independently classify $n$ subjects within $K$ categories, it is common to use different kappa type coefficients, the most common of which is the $\kappa_C$ coefficient (Cohen's…

统计理论 · 数学 2026-02-24 A. Martín Andrés , M. Álvarez Hernández

Cohen's and Fleiss' kappa are well-known measures of inter-rater agreement, but they restrict each rater to selecting only one category per subject. This limitation is consequential in contexts where subjects may belong to multiple…

统计方法学 · 统计学 2025-09-22 Filip Moons , Ellen Vandervieren

The state of the art in human computer conversation leaves something to be desired and, indeed, talking to a computer can be down-right annoying. This paper describes an approach to identifying ``opportunities for improvement'' in these…

计算与语言 · 计算机科学 2023-07-17 Peter Wallis

A novel approach for comparing quality attributes of different products when there is considerable product-related variability is proposed. In such a case, the whole range of possible realizations must be considered. Looking, for example,…

统计方法学 · 统计学 2024-08-30 Gerhard Gössler , Vera Hofer , Hans Manner , Walter Goessler

Complex assignments typically consist of open-ended questions with large and diverse content in the context of both classroom and online graduate programs. With the sheer scale of these programs comes a variety of problems in peer and…

计算与语言 · 计算机科学 2020-03-17 Manikandan Ravikiran

Inter-coder agreement measures, like Cohen's kappa, correct the relative frequency of agreement between coders to account for agreement which simply occurs by chance. However, in some situations these measures exhibit behavior which make…

应用统计 · 统计学 2012-08-07 Dirk Schuster

How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric $\kappa_p$ applied to 20 languages and 47 subjects in GlobalMMLU. Our analysis reveals that a model's…

计算与语言 · 计算机科学 2025-10-03 Debangan Mishra , Arihant Rastogi , Agyeya Negi , Shashwat Goel , Ponnurangam Kumaraguru

We introduce LAMBADA, a dataset to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative passages sharing the characteristic that human subjects are…

Discourse signals are often implicit, leaving it up to the interpreter to draw the required inferences. At the same time, discourse is embedded in a social context, meaning that interpreters apply their own assumptions and beliefs when…

计算与语言 · 计算机科学 2021-04-12 Elisa Ferracane , Greg Durrett , Junyi Jessy Li , Katrin Erk

Textbooks on statistics emphasize care and precision, via concepts such as reliability and validity in measurement, random sampling and treatment assignment in data collection, and causal identification and bias in estimation. But how do…

统计方法学 · 统计学 2013-07-24 Andrew Gelman , Keith O'Rourke

Text-based conversational agents (CAs) are increasingly used in mental health, yet evaluation practices remain fragmented. We conducted a PRISMA-guided systematic review (May-June 2024) across ACM Digital Library, Scopus, and PsycINFO. From…

人机交互 · 计算机科学 2026-02-23 Jiangtao Gong , Xiao Wen , Fengyi Tao , Xinqi Wang , Xixi Yang , Yangrong Tang

In this paper, we propose standard statistical tools as a solution to commonly highlighted problems in the explainability literature. Indeed, leveraging statistical estimators allows for a proper definition of explanations, enabling…

机器学习 · 统计学 2024-05-01 Valentina Ghidini

Determining the plausibility of causal relations between clauses is a commonsense reasoning task that requires complex inference ability. The general approach to this task is to train a large pretrained language model on a specific dataset.…

计算与语言 · 计算机科学 2021-01-14 Ieva Staliūnaitė , Philip John Gorinski , Ignacio Iacobacci

Online discourse is often perceived as polarized and unproductive. While some conversational discourse parsing frameworks are available, they do not naturally lend themselves to the analysis of contentious and polarizing discussions.…

计算与语言 · 计算机科学 2020-12-09 Stepan Zakharov , Omri Hadar , Tovit Hakak , Dina Grossman , Yifat Ben-David Kolikant , Oren Tsur

Language is a social phenomenon and variation is inherent to its social nature. Recently, there has been a surge of interest within the computational linguistics (CL) community in the social dimension of language. In this article we present…

计算与语言 · 计算机科学 2016-04-07 Dong Nguyen , A. Seza Doğruöz , Carolyn P. Rosé , Franciska de Jong

Heuristics and cognitive biases are an integral part of human decision-making. Automatically detecting a particular cognitive bias could enable intelligent tools to provide better decision-support. Detecting the presence of a cognitive bias…

人机交互 · 计算机科学 2024-01-15 Stephen Pilli

A discourse strategy is a strategy for communicating with another agent. Designing effective dialogue systems requires designing agents that can choose among discourse strategies. We claim that the design of effective strategies must take…

cmp-lg · 计算机科学 2016-08-31 Marilyn A. Walker

Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers. Automatic metrics such as BLEU correlate weakly with human annotations, resulting in a significant bias across different models and…

计算与语言 · 计算机科学 2020-04-02 Nouha Dziri , Ehsan Kamalloo , Kory W. Mathewson , Osmar Zaiane

During the past decade, several areas of speech and language understanding have witnessed substantial breakthroughs from the use of data-driven models. In the area of dialogue systems, the trend is less obvious, and most practical systems…

计算与语言 · 计算机科学 2017-03-22 Iulian Vlad Serban , Ryan Lowe , Peter Henderson , Laurent Charlin , Joelle Pineau
‹ 上一页 1 2 3 10 下一页 ›