中文
相关论文

相关论文: The Ecological Fallacy in Annotation: Modelling Hu…

200 篇论文

People naturally vary in their annotations for subjective questions and some of this variation is thought to be due to the person's sociodemographic characteristics. LLMs have also been used to label data, but recent work has shown that…

计算与语言 · 计算机科学 2025-03-03 Matthias Orlikowski , Jiaxin Pei , Paul Röttger , Philipp Cimiano , David Jurgens , Dirk Hovy

When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to predicting individual annotator ratings on the toxicity of text…

计算与语言 · 计算机科学 2024-10-17 Harbani Jaggi , Kashyap Murali , Eve Fleisig , Erdem Bıyık

Human label variation has been established as a central phenomenon in NLP: the perspectives different annotators have on the same item need to be embraced. Data collection practices thus shifted towards increasing the annotator numbers and…

计算与语言 · 计算机科学 2026-05-08 Maximilian Maurer , Maximilian Linde , Gabriella Lapesa

Understanding the sources of variability in annotations is crucial for developing fair NLP systems, especially for tasks like sexism detection where demographic bias is a concern. This study investigates the extent to which annotator…

计算与语言 · 计算机科学 2025-07-29 Hadi Mohammadi , Tina Shahedi , Pablo Mosteiro , Massimo Poesio , Ayoub Bagheri , Anastasia Giachanou

A common practice in building NLP datasets, especially using crowd-sourced annotations, involves obtaining multiple annotator judgements on the same data instances, which are then flattened to produce a single "ground truth" label or score,…

计算与语言 · 计算机科学 2021-10-13 Vinodkumar Prabhakaran , Aida Mostafazadeh Davani , Mark Díaz

Demographics and cultural background of annotators influence the labels they assign in text annotation -- for instance, an elderly woman might find it offensive to read a message addressed to a "bro", but a male teenager might find it…

Human label variation, or annotation disagreement, exists in many natural language processing (NLP) tasks, including natural language inference (NLI). To gain direct evidence of how NLI label variation arises, we build LiveNLI, an English…

计算与语言 · 计算机科学 2023-10-24 Nan-Jiang Jiang , Chenhao Tan , Marie-Catherine de Marneffe

In NLP annotation, it is common to have multiple annotators label the text and then obtain the ground truth labels based on the agreement of major annotators. However, annotators are individuals with different backgrounds, and minors'…

计算与语言 · 计算机科学 2023-01-13 Ruyuan Wan , Jaehyung Kim , Dongyeop Kang

Human label variation (Plank 2022), or annotation disagreement, exists in many natural language processing (NLP) tasks. To be robust and trusted, NLP models need to identify such variation and be able to explain it. To this end, we created…

计算与语言 · 计算机科学 2023-04-26 Nan-Jiang Jiang , Chenhao Tan , Marie-Catherine de Marneffe

Humans often hold different perspectives on the same issues. In many NLP tasks, annotation disagreement can reflect valid subjective perspectives. Modeling annotator perspectives and understanding their relationship with other human…

计算与语言 · 计算机科学 2026-04-21 Leixin Zhang , Cagri Coltekin

Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large language models (LLMs) can learn and reproduce such…

计算与语言 · 计算机科学 2026-05-28 Beiduo Chen , Pingjun Hong , Ziyun Zhang , Benjamin Roth , Anna Korhonen , Barbara Plank

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic…

Data annotation remains the sine qua non of machine learning and AI. Recent empirical work on data annotation has begun to highlight the importance of rater diversity for fairness, model performance, and new lines of research have begun to…

人工智能 · 计算机科学 2024-02-13 Andrew Smart , Ding Wang , Ellis Monk , Mark Díaz , Atoosa Kasirzadeh , Erin Van Liemt , Sonja Schmer-Galunder

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them massively generate…

计算与语言 · 计算机科学 2019-08-29 Mor Geva , Yoav Goldberg , Jonathan Berant

We commonly use agreement measures to assess the utility of judgements made by human annotators in Natural Language Processing (NLP) tasks. While inter-annotator agreement is frequently used as an indication of label reliability by…

计算与语言 · 计算机科学 2025-10-21 Gavin Abercrombie , Tanvi Dinkar , Amanda Cercas Curry , Verena Rieser , Dirk Hovy

The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and…

计算与语言 · 计算机科学 2022-05-11 Maarten Sap , Swabha Swayamdipta , Laura Vianna , Xuhui Zhou , Yejin Choi , Noah A. Smith

Researchers have proposed the use of generative large language models (LLMs) to label data for research and applied settings. This literature emphasizes the improved performance of these models relative to other natural language models,…

计算与语言 · 计算机科学 2025-06-17 Megan A. Brown , Shubham Atreja , Libby Hemphill , Patrick Y. Wu

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated…

计算与语言 · 计算机科学 2024-05-17 Negar Mokhberian , Myrl G. Marmarelis , Frederic R. Hopp , Valerio Basile , Fred Morstatter , Kristina Lerman

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

计算与语言 · 计算机科学 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

The development of real-time affect detection models often depends upon obtaining annotated data for supervised learning by employing human experts to label the student data. One open question in annotating affective data for affect…

人机交互 · 计算机科学 2019-01-15 Eda Okur , Sinem Aslan , Nese Alyuz , Asli Arslan Esme , Ryan S. Baker
‹ 上一页 1 2 3 10 下一页 ›