中文
相关论文

相关论文: Quantifying and Predicting Disagreement in Graded …

200 篇论文

Researchers have raised awareness about the harms of aggregating labels especially in subjective tasks that naturally contain disagreements among human annotators. In this work we show that models that are only provided aggregated labels…

This paper investigates the role of text in visualizations, specifically the impact of text position, semantic content, and biased wording. Two empirical studies were conducted based on two tasks (predicting data trends and appraising bias)…

人机交互 · 计算机科学 2024-01-09 Chase Stokes , Cindy Xiong Bearfield , Marti A. Hearst

In this paper, we investigate how personalising Large Language Models (Persona-LLMs) with annotator personas affects their sensitivity to hate speech, particularly regarding biases linked to shared or differing identities between annotators…

计算与语言 · 计算机科学 2025-10-23 Ewelina Gajewska , Arda Derbent , Jaroslaw A Chudziak , Katarzyna Budzynska

Large Language Models (LLMs) have shown strong performance on NLP classification tasks. However, they typically rely on aggregated labels-often via majority voting-which can obscure the human disagreement inherent in subjective annotations.…

计算与语言 · 计算机科学 2025-06-09 Benedetta Muscato , Yue Li , Gizem Gezici , Zhixue Zhao , Fosca Giannotti

Subjectivity and difference of opinion are key social phenomena, and it is crucial to take these into account in the annotation and detection process of derogatory textual content. In this paper, we use four datasets provided by…

计算与语言 · 计算机科学 2023-05-03 Sadat Shahriar , Thamar Solorio

We propose a novel method to conceptually decompose an existing annotation into separate levels, allowing the analysis of inter-annotators disagreement in each level separately. We suggest two distinct strategies in order to actualize this…

计算与语言 · 计算机科学 2025-06-11 Effi Levi , Shaul R. Shenhav

Assigning a positive or negative score to a word out of context (i.e. a word's prior polarity) is a challenging task for sentiment analysis. In the literature, various approaches based on SentiWordNet have been proposed. In this paper, we…

计算与语言 · 计算机科学 2013-09-24 Marco Guerini , Lorenzo Gatti , Marco Turchi

The annotation of textual information is a fundamental activity in Linguistics and Computational Linguistics. This article presents various observations on annotations. It approaches the topic from several angles including Hypertext,…

计算与语言 · 计算机科学 2020-04-23 Georg Rehm

Do LLMs align with human perceptions of safety? We study this question via annotation alignment, the extent to which LLMs and humans agree when annotating the safety of user-chatbot conversations. We leverage the recent DICES dataset (Aroyo…

计算与语言 · 计算机科学 2024-10-08 Rajiv Movva , Pang Wei Koh , Emma Pierson

Computational humor detection systems rarely model the subjectivity of humor responses, or consider alternative reactions to humor - namely offense. We analyzed a large dataset of humor and offense ratings by male and female annotators of…

计算与语言 · 计算机科学 2022-08-24 J. A. Meaney , Steven R. Wilson , Luis Chiruzzo , Walid Magdy

Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for Area Chairs and editors to reliably identify and interpret such disagreements. Existing…

计算与语言 · 计算机科学 2026-05-12 Sandeep Kumar , Yash Kamdar , Abid Hossain , Bharti Kumari , Tanik Saikh , Asif Ekbal

Hate speech moderation remains a challenging task for social media platforms. Human-AI collaborative systems offer the potential to combine the strengths of humans' reliability and the scalability of machine learning to tackle this issue…

人机交互 · 计算机科学 2023-07-25 Philippe Lammerts , Philip Lippmann , Yen-Chia Hsu , Fabio Casati , Jie Yang

In this paper, we present findings from an semi-experimental exploration of rater diversity and its influence on safety annotations of conversations generated by humans talking to a generative AI-chat bot. We find significant differences in…

人机交互 · 计算机科学 2023-05-12 Lora Aroyo , Mark Diaz , Christopher Homan , Vinodkumar Prabhakaran , Alex Taylor , Ding Wang

Prior work has revealed that positive words occur more frequently than negative words in human expressions, which is typically attributed to positivity bias, a tendency for people to report positive views of reality. But what about the…

计算与语言 · 计算机科学 2021-06-24 Madhusudhan Aithal , Chenhao Tan

We report here on a study of interannotator agreement in the coreference task as defined by the Message Understanding Conference (MUC-6 and MUC-7). Based on feedback from annotators, we clarified and simplified the annotation specification.…

cmp-lg · 计算机科学 2007-05-23 Lynette Hirschman , Patricia Robinson , John Burger , Marc Vilain

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model…

计算与语言 · 计算机科学 2025-03-03 Matthias Orlikowski , Paul Röttger , Philipp Cimiano , Dirk Hovy

Stance detection, which aims to determine whether an individual is for or against a target concept, promises to uncover public opinion from large streams of social media data. Yet even human annotation of social media content does not…

社会与信息网络 · 计算机科学 2021-09-08 Kenneth Joseph , Sarah Shugars , Ryan Gallagher , Jon Green , Alexi Quintana Mathé , Zijian An , David Lazer

Sentiment analysis is an important tool for aggregating patient voices, in order to provide targeted improvements in healthcare services. A prerequisite for this is the availability of in-domain data annotated for sentiment. This article…

Sentiment analysis AKA opinion mining is one of the most widely used NLP applications to identify human intentions from their reviews. In the education sector, opinion mining is used to listen to student opinions and enhance their…

计算与语言 · 计算机科学 2023-02-10 Thanveer Shaik , Xiaohui Tao , Christopher Dann , Haoran Xie , Yan Li , Linda Galligan

This paper presents a simple unsupervised learning algorithm for classifying reviews as recommended (thumbs up) or not recommended (thumbs down). The classification of a review is predicted by the average semantic orientation of the phrases…

机器学习 · 计算机科学 2007-05-23 Peter D. Turney
‹ 上一页 1 8 9 10 下一页 ›