中文
相关论文

相关论文: Re-examining Sexism and Misogyny Classification wi…

200 篇论文

Resolving disagreement in manual annotation typically consists of removing unreliable annotators and using a label aggregation strategy such as majority vote or expert opinion to resolve disagreement. These may have the side-effect of…

计算与语言 · 计算机科学 2024-12-06 Mugdha Pandya , Nafise Sadat Moosavi , Diana Maynard

Content moderation and toxicity classification represent critical tasks with significant social implications. However, studies have shown that major classification models exhibit tendencies to magnify or reduce biases and potentially…

Crowdsourced annotation is vital to both collecting labelled data to train and test automated content moderation systems and to support human-in-the-loop review of system decisions. However, annotation tasks such as judging hate speech are…

Identifying misogyny using artificial intelligence is a form of combating online toxicity against women. However, the subjective nature of interpreting misogyny poses a significant challenge to model the phenomenon. In this paper, we…

计算与语言 · 计算机科学 2024-06-25 Jason Angel , Segun Taofeek Aroyehun , Grigori Sidorov , Alexander Gelbukh

Language serves as a powerful tool for the manifestation of societal belief systems. In doing so, it also perpetuates the prevalent biases in our society. Gender bias is one of the most pervasive biases in our society and is seen in online…

计算与语言 · 计算机科学 2023-10-27 Rishav Hada , Agrima Seth , Harshita Diddee , Kalika Bali

Public institutions are increasingly reliant on data from social media sites to measure public attitude and provide timely public engagement. Such reliance includes the exploration of public views on important social issues such as…

社会与信息网络 · 计算机科学 2015-06-30 Hemant Purohit , Tanvi Banerjee , Andrew Hampton , Valerie L. Shalin , Nayanesh Bhandutia , Amit P. Sheth

Classifiers tend to propagate biases present in the data on which they are trained. Hence, it is important to understand how the demographic identities of the annotators of comments affect the fairness of the resulting model. In this paper,…

计算与语言 · 计算机科学 2021-06-07 Elizabeth Excell , Noura Al Moubayed

It is common practice in text classification to only use one majority label for model training even if a dataset has been annotated by multiple annotators. Doing so can remove valuable nuances and diverse perspectives inherent in the…

计算与语言 · 计算机科学 2024-09-27 Jin Xu , Mariët Theune , Daniel Braun

Different linguistic expressions can conceptualize the same event from different viewpoints by emphasizing certain participants over others. Here, we investigate a case where this has social consequences: how do linguistic expressions of…

计算与语言 · 计算机科学 2022-09-27 Gosse Minnema , Sara Gemelli , Chiara Zanchi , Tommaso Caselli , Malvina Nissim

Gender-based violence (GBV) is a major public health issue, with the World Health Organization estimating that one in three women experiences physical or sexual violence by an intimate partner during her lifetime. In Brazil, although…

Though majority vote among annotators is typically used for ground truth labels in natural language processing, annotator disagreement in tasks such as hate speech detection may reflect differences in opinion across groups, not noise. Thus,…

计算与语言 · 计算机科学 2024-03-19 Eve Fleisig , Rediet Abebe , Dan Klein

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic…

Since state-of-the-art approaches to offensive language detection rely on supervised learning, it is crucial to quickly adapt them to the continuously evolving scenario of social media. While several approaches have been proposed to tackle…

计算与语言 · 计算机科学 2022-10-17 Elisa Leonardelli , Stefano Menini , Alessio Palmero Aprosio , Marco Guerini , Sara Tonelli

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

计算与语言 · 计算机科学 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated…

计算与语言 · 计算机科学 2024-05-17 Negar Mokhberian , Myrl G. Marmarelis , Frederic R. Hopp , Valerio Basile , Fred Morstatter , Kristina Lerman

Different ways of linguistically expressing the same real-world event can lead to different perceptions of what happened. Previous work has shown that different descriptions of gender-based violence (GBV) influence the reader's perception…

计算与语言 · 计算机科学 2023-06-02 Gosse Minnema , Huiyuan Lai , Benedetta Muscato , Malvina Nissim

Majority voting and averaging are common approaches employed to resolve annotator disagreements and derive single ground truth labels from multiple annotations. However, annotators may systematically disagree with one another, often…

计算与语言 · 计算机科学 2021-10-13 Aida Mostafazadeh Davani , Mark Díaz , Vinodkumar Prabhakaran

Data annotation, the practice of assigning descriptive labels to raw data, is pivotal in optimizing the performance of machine learning models. However, it is a resource-intensive process susceptible to biases introduced by annotators. The…

Gender-based violence (GBV) is a human-generated crisis, existing in various forms, including offline, via physical and sexual violence, and now online via harassment and trolling. While studying social media campaigns for different domains…

社会与信息网络 · 计算机科学 2016-08-05 Prakruthi Karuna , Hemant Purohit , Bonnie Stabile , Angela Hattery

The widespread adoption of automatic sentiment and emotion classifiers makes it important to ensure that these tools perform reliably across different populations. Yet their reliability is typically assessed using benchmarks that rely on…

计算与语言 · 计算机科学 2026-01-09 Ivan Smirnov , Segun T. Aroyehun , Paul Plener , David Garcia
‹ 上一页 1 2 3 10 下一页 ›