中文
相关论文

相关论文: NLPositionality: Characterizing Design Biases of D…

200 篇论文

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them massively generate…

计算与语言 · 计算机科学 2019-08-29 Mor Geva , Yoav Goldberg , Jonathan Berant

Annotation bias in NLP datasets remains a major challenge for developing multilingual Large Language Models (LLMs), particularly in culturally diverse settings. Bias from task framing, annotator subjectivity, and cultural mismatches can…

计算与语言 · 计算机科学 2025-11-19 Xia Cui , Ziyi Huang , Naeemeh Adel

Subjective judgments are part of several NLP datasets and recent work is increasingly prioritizing models whose outputs reflect this diversity of perspectives. Such responses allow us to shed light on minority voices, which are frequently…

计算与语言 · 计算机科学 2026-03-31 Urja Khurana , Michiel van der Meer , Enrico Liscio , Antske Fokkens , Pradeep K. Murukannaiah

Using novel approaches to dataset development, the Biasly dataset captures the nuance and subtlety of misogyny in ways that are unique within the literature. Built in collaboration with multi-disciplinary experts and annotators themselves,…

This paper is a summary of the work done in my PhD thesis. Where I investigate the impact of bias in NLP models on the task of hate speech detection from three perspectives: explainability, offensive stereotyping bias, and fairness. Then, I…

计算与语言 · 计算机科学 2023-12-06 Fatma Elsafoury

We test whether NLP datasets created with Large Language Models (LLMs) contain annotation artifacts and social biases like NLP datasets elicited from crowd-source workers. We recreate a portion of the Stanford Natural Language Inference…

计算与语言 · 计算机科学 2025-03-10 Grace Proebsting , Adam Poliak

Natural Language Inference (NLI) is foundational for evaluating language understanding in AI. However, progress has plateaued, with models failing on ambiguous examples and exhibiting poor generalization. We argue that this stems from…

计算与语言 · 计算机科学 2024-05-21 Claudiu Creanga , Liviu P. Dinu

While human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection. Recent studies that have critically examined this issue are often situated in the Western context, and…

计算与语言 · 计算机科学 2024-04-18 Aida Mostafazadeh Davani , Mark Díaz , Dylan Baker , Vinodkumar Prabhakaran

Language models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias. A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy…

计算与语言 · 计算机科学 2023-07-07 Shangbin Feng , Chan Young Park , Yuhan Liu , Yulia Tsvetkov

Annotator disagreement is widespread in NLP, particularly for subjective and ambiguous tasks such as toxicity detection and stance analysis. While early approaches treated disagreement as noise to be removed, recent work increasingly models…

计算与语言 · 计算机科学 2026-01-21 Yinuo Xu , David Jurgens

A common practice in building NLP datasets, especially using crowd-sourced annotations, involves obtaining multiple annotator judgements on the same data instances, which are then flattened to produce a single "ground truth" label or score,…

计算与语言 · 计算机科学 2021-10-13 Vinodkumar Prabhakaran , Aida Mostafazadeh Davani , Mark Díaz

Annotators' sociodemographic backgrounds (i.e., the individual compositions of their gender, age, educational background, etc.) have a strong impact on their decisions when working on subjective NLP tasks, such as toxic language detection.…

计算与语言 · 计算机科学 2024-02-09 Tilman Beck , Hendrik Schuff , Anne Lauscher , Iryna Gurevych

Language data and models demonstrate various types of bias, be it ethnic, religious, gender, or socioeconomic. AI/NLP models, when trained on the racially biased dataset, AI/NLP models instigate poor model explainability, influence user…

计算与语言 · 计算机科学 2022-11-28 Kinshuk Sengupta , Praveen Ranjan Srivastava

Modern models for common NLP tasks often employ machine learning techniques and train on journalistic, social media, or other culturally-derived text. These have recently been scrutinized for racial and gender biases, rooting from inherent…

计算与语言 · 计算机科学 2026-01-27 Scott Friedman , Sonja Schmer-Galunder , Anthony Chen , Jeffrey Rye

We investigate the potential for nationality biases in natural language processing (NLP) models using human evaluation methods. Biased NLP models can perpetuate stereotypes and lead to algorithmic discrimination, posing a significant…

Data-driven statistical Natural Language Processing (NLP) techniques leverage large amounts of language data to build models that can understand language. However, most language data reflect the public discourse at the time the data was…

计算与语言 · 计算机科学 2019-10-11 Vinodkumar Prabhakaran , Ben Hutchinson , Margaret Mitchell

With the rapid proliferation of artificial intelligence, there is growing concern over its potential to exacerbate existing biases and societal disparities and introduce novel ones. This issue has prompted widespread attention from…

人机交互 · 计算机科学 2024-05-01 Sanjana Gautam , Mukund Srinath

The automatic detection of hate speech online is an active research area in NLP. Most of the studies to date are based on social media datasets that contribute to the creation of hate speech detection models trained on them. However, data…

计算与语言 · 计算机科学 2023-07-06 Dimosthenis Antypas , Jose Camacho-Collados

Building equitable and inclusive NLP technologies demands consideration of whether and how social attitudes are represented in ML models. In particular, representations encoded in models often inadvertently perpetuate undesirable social…

计算与语言 · 计算机科学 2020-05-05 Ben Hutchinson , Vinodkumar Prabhakaran , Emily Denton , Kellie Webster , Yu Zhong , Stephen Denuyl

Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-stakes settings affecting autistic communities. While previous research has identified…

计算与语言 · 计算机科学 2026-05-27 Naba Rizvi , Harper Strickland , Saleha Ahmedi , Nedjma Ousidhoum
‹ 上一页 1 2 3 10 下一页 ›