English
Related papers

Related papers: GeniL: A Multilingual Dataset on Generalizing Lang…

200 papers

With the ever-increasing cases of hate spread on social media platforms, it is critical to design abuse detection mechanisms to proactively avoid and control such incidents. While there exist methods for hate speech detection, they…

Computation and Language · Computer Science 2020-01-17 Pinkesh Badjatiya , Manish Gupta , Vasudeva Varma

Sentiment analysis is a common task in natural language processing that aims to detect polarity of a text document (typically a consumer review). In the simplest settings, we discriminate only between positive and negative sentiment,…

Computation and Language · Computer Science 2015-05-28 Grégoire Mesnil , Tomas Mikolov , Marc'Aurelio Ranzato , Yoshua Bengio

Large language models (LLMs) are trained on vast, uncurated datasets that contain various forms of biases and language reinforcing harmful stereotypes that may be subsequently inherited by the models themselves. Therefore, it is essential…

Computation and Language · Computer Science 2024-10-01 Jacob-Junqi Tian , Omkar Dige , D. B. Emerson , Faiza Khan Khattak

Stereotypes influence social perceptions and can escalate into discrimination and violence. While NLP research has extensively addressed gender bias and hate speech, stereotype detection remains an emerging field with significant societal…

Computation and Language · Computer Science 2025-10-08 Alessandra Teresa Cignarella , Anastasia Giachanou , Els Lefever

Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29…

Computation and Language · Computer Science 2026-02-24 Jacqueline Rowe , Mateusz Klimaszewski , Liane Guillou , Shannon Vallor , Alexandra Birch

Recent studies have shown that generative language models often reflect and amplify societal biases in their outputs. However, these studies frequently conflate observed biases with other task-specific shortcomings, such as comprehension…

Computation and Language · Computer Science 2024-12-17 Akshita Jha , Sanchit Kabra , Chandan K. Reddy

Lack of repeatability and generalisability are two significant threats to continuing scientific development in Natural Language Processing. Language models and learning methods are so complex that scientific conference papers no longer…

Computation and Language · Computer Science 2018-08-07 Andrew Moore , Paul Rayson

Gender bias represents a form of systematic negative treatment that targets individuals based on their gender. This discrimination can range from subtle sexist remarks and gendered stereotypes to outright hate speech. Prior research has…

Computation and Language · Computer Science 2024-03-19 Karolina Stańczak

The rapid deployment of generative language models (LMs) has raised concerns about social biases affecting the well-being of diverse consumers. The extant literature on generative LMs has primarily examined bias via explicit identity…

Computation and Language · Computer Science 2026-05-04 Evan Shieh , Faye-Marie Vassel , Cassidy Sugimoto , Thema Monroe-White

Stereotype detection is a challenging and subjective task, as certain statements, such as "Black people like to play basketball," may not appear overtly toxic but still reinforce racial stereotypes. With the increasing prevalence of large…

Computation and Language · Computer Science 2024-11-19 Zekun Wu , Sahan Bulathwela , Maria Perez-Ortiz , Adriano Soares Koshiyama

Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data. This has been linked to dataset biases and the use of sentence-level labels, which fail to teach models the…

Computation and Language · Computer Science 2025-06-05 Agostina Calabrese , Tom Sherborne , Björn Ross , Mirella Lapata

Recent works have found evidence of gender bias in models of machine translation and coreference resolution using mostly synthetic diagnostic datasets. While these quantify bias in a controlled experiment, they often do so on a small scale…

Computation and Language · Computer Science 2021-09-13 Shahar Levy , Koren Lazar , Gabriel Stanovsky

A stereotype is an over-generalized belief about a particular group of people, e.g., Asians are good at math or Asians are bad drivers. Such beliefs (biases) are known to hurt target groups. Since pretrained language models are trained on…

Computation and Language · Computer Science 2020-04-21 Moin Nadeem , Anna Bethke , Siva Reddy

We present a methodological framework to discover linguistic and discursive patterns associated to different social groups through contrastive synthetic text generation and statistical analysis. In contrast with previous approaches, we aim…

Computation and Language · Computer Science 2026-04-21 S. A. Desimone , L. Alonso Alemany

We present GEST -- a new manually created dataset designed to measure gender-stereotypical reasoning in language models and machine translation systems. GEST contains samples for 16 gender stereotypes about men and women (e.g., Women are…

Computation and Language · Computer Science 2024-10-02 Matúš Pikuliak , Andrea Hrckova , Stefan Oresko , Marián Šimko

The design of Large Language Models and generative artificial intelligence has been shown to be "unfair" to less-spoken languages and to deepen the digital language divide. Critical sociolinguistic work has also argued that these…

Computation and Language · Computer Science 2026-03-31 Verena Platzgummer , John McCrae , Sina Ahmadi

Social categories and stereotypes are embedded in language and can introduce data bias into Large Language Models (LLMs). Despite safeguards, these biases often persist in model behavior, potentially leading to representational harm in…

Computation and Language · Computer Science 2025-02-27 Rebekka Görge , Michael Mock , Héctor Allende-Cid

Meta-learning has been shown to have better performance than supervised learning for few-shot monolingual spoken word classification. However, the meta-learning approach remains under-explored in multilingual spoken word classification. In…

Computation and Language · Computer Science 2026-05-15 Batsirayi Mupamhi Ziki , Louise Beyers , Ruan van der Merwe

The prevalence of offensive content on the internet, encompassing hate speech and cyberbullying, is a pervasive issue worldwide. Consequently, it has garnered significant attention from the machine learning (ML) and natural language…

Computation and Language · Computer Science 2024-07-29 Alphaeus Dmonte , Tejas Arya , Tharindu Ranasinghe , Marcos Zampieri

A significant portion of the textual data used in the field of Natural Language Processing (NLP) exhibits gender biases, particularly due to the use of masculine generics (masculine words that are supposed to refer to mixed groups of men…

Computation and Language · Computer Science 2025-05-30 Enzo Doyen , Amalia Todirascu