中文
相关论文

相关论文: When Does Demographic Information Help? Data and M…

200 篇论文

Supervised approaches generally rely on majority-based labels. However, it is hard to achieve high agreement among annotators in subjective tasks such as hate speech detection. Existing neural network models principally regard labels as…

计算与语言 · 计算机科学 2023-01-11 Wenjie Yin , Vibhor Agarwal , Aiqi Jiang , Arkaitz Zubiaga , Nishanth Sastry

When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by annotators' social identities and lived experiences. Yet standard practice still flattens…

人工智能 · 计算机科学 2026-04-10 Samay U. Shetty , Tharindu Cyril Weerasooriya , Deepak Pandita , Christopher M. Homan

When training data are collected from human annotators, the design of the annotation instrument, the instructions given to annotators, the characteristics of the annotators, and their interactions can impact training data. This study…

机器学习 · 统计学 2024-01-23 Christoph Kern , Stephanie Eckman , Jacob Beck , Rob Chew , Bolei Ma , Frauke Kreuter

Voice-based interfaces are widely used; however, achieving fair Wake-up Word detection across diverse speaker populations remains a critical challenge due to persistent demographic biases. This study evaluates the effectiveness of…

计算与语言 · 计算机科学 2026-04-08 Fernando López , Paula Delgado-Santos , Pablo Gómez , David Solans , Jordi Luque

Annotators' sociodemographic backgrounds (i.e., the individual compositions of their gender, age, educational background, etc.) have a strong impact on their decisions when working on subjective NLP tasks, such as toxic language detection.…

计算与语言 · 计算机科学 2024-02-09 Tilman Beck , Hendrik Schuff , Anne Lauscher , Iryna Gurevych

In this work, we explore the capability of Large Language Models (LLMs) to annotate hate speech and abusiveness while considering predefined annotator personas within the strong-to-weak data perspectivism spectra. We evaluated LLM-generated…

计算与语言 · 计算机科学 2025-08-26 Olufunke O. Sarumi , Charles Welch , Daniel Braun , Jörg Schlötterer

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

计算与语言 · 计算机科学 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

Demographic factors (e.g., gender or age) shape our language. Previous work showed that incorporating demographic factors can consistently improve performance for various NLP tasks with traditional NLP models. In this work, we investigate…

计算与语言 · 计算机科学 2023-05-10 Chia-Chien Hung , Anne Lauscher , Dirk Hovy , Simone Paolo Ponzetto , Goran Glavaš

Algorithms deployed in education can shape the learning experience and success of a student. It is therefore important to understand whether and how such algorithms might create inequalities or amplify existing biases. In this paper, we…

计算机与社会 · 计算机科学 2022-12-21 Jade Maï Cock , Muhammad Bilal , Richard Davis , Mirko Marras , Tanja Käser

Understanding the sources of variability in annotations is crucial for developing fair NLP systems, especially for tasks like sexism detection where demographic bias is a concern. This study investigates the extent to which annotator…

计算与语言 · 计算机科学 2025-07-29 Hadi Mohammadi , Tina Shahedi , Pablo Mosteiro , Massimo Poesio , Ayoub Bagheri , Anastasia Giachanou

This paper investigates how hate speech varies in systematic ways according to the identities it targets. Across multiple hate speech datasets annotated for targeted identities, we find that classifiers trained on hate speech targeting…

计算与语言 · 计算机科学 2022-12-08 Michael Miller Yoder , Lynnette Hui Xian Ng , David West Brown , Kathleen M. Carley

In many domains, it is difficult to obtain the race data that is required to estimate racial disparity. To address this problem, practitioners have adopted the use of proxy methods which predict race using non-protected covariates. However,…

计算机与社会 · 计算机科学 2024-09-04 Kweku Kwegyir-Aggrey , Naveen Durvasula , Jennifer Wang , Suresh Venkatasubramanian

Unwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets. In this work, we ask whether training on demographically perturbed data leads to fairer language models. We collect…

计算与语言 · 计算机科学 2022-10-14 Rebecca Qian , Candace Ross , Jude Fernandes , Eric Smith , Douwe Kiela , Adina Williams

Fairness in human-robot interaction critically depends on the reliability of the perceptual models that enable robots to interpret human behavior. While demographic biases have been widely studied in high-level facial analysis tasks, their…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Pablo Parte , Roberto Valle , José M. Buenaposada , Luis Baumela

Statistical agencies rely on sampling techniques to collect socio-demographic data crucial for policy-making and resource allocation. This paper shows that surveys of important societal relevance introduce sampling errors that unevenly…

密码学与安全 · 计算机科学 2025-01-22 Joonhyuk Ko , Juba Ziani , Saswat Das , Matt Williams , Ferdinando Fioretto

Human biases have been shown to influence the performance of models and algorithms in various fields, including Natural Language Processing. While the study of this phenomenon is garnering focus in recent years, the available resources are…

计算与语言 · 计算机科学 2024-08-15 Ana Sofia Evans , Helena Moniz , Luísa Coheur

Estimating population parameters in finite populations of text documents can be challenging when obtaining the labels for the target variable requires manual annotation. To address this problem, we combine predictions from a transformer…

计算与语言 · 计算机科学 2025-05-09 Hannes Waldetoft , Jakob Torgander , Måns Magnusson

Building a benchmark dataset for hate speech detection presents various challenges. Firstly, because hate speech is relatively rare, random sampling of tweets to annotate is very inefficient in finding hate speech. To address this, prior…

计算与语言 · 计算机科学 2021-11-11 Md Mustafizur Rahman , Dinesh Balakrishnan , Dhiraj Murthy , Mucahid Kutlu , Matthew Lease

Incorporating every annotator's perspective is crucial for unbiased data modeling. Annotator fatigue and changing opinions over time can distort dataset annotations. To combat this, we propose to learn a more accurate representation of…

机器学习 · 计算机科学 2024-06-05 Uthman Jinadu , Yi Ding

Recent work introduced the model of learning from discriminative feature feedback, in which a human annotator not only provides labels of instances, but also identifies discriminative features that highlight important differences between…

机器学习 · 计算机科学 2021-05-25 Sanjoy Dasgupta , Sivan Sabato