中文
相关论文

相关论文: Combating high variance in Data-Scarce Implicit Ha…

200 篇论文

Hate speech and profanity detection suffer from data sparsity, especially for languages other than English, due to the subjective nature of the tasks and the resulting annotation incompatibility of existing corpora. In this study, we…

计算与语言 · 计算机科学 2021-06-21 Vanessa Hahn , Dana Ruiter , Thomas Kleinbauer , Dietrich Klakow

Detection of some types of toxic language is hampered by extreme scarcity of labeled training data. Data augmentation - generating new synthetic data from a labeled seed dataset - can help. The efficacy of data augmentation on toxic…

计算与语言 · 计算机科学 2020-10-27 Mika Juuti , Tommi Gröndahl , Adrian Flanagan , N. Asokan

In current hate speech datasets, there exists a high correlation between annotators' perceptions of toxicity and signals of African American English (AAE). This bias in annotated training data and the tendency of machine learning models to…

计算与语言 · 计算机科学 2020-05-26 Mengzhou Xia , Anjalie Field , Yulia Tsvetkov

This paper addresses the critical challenge of developing computationally efficient hate speech detection systems that maintain competitive performance while being practical for real-time deployment. We propose a novel three-layer framework…

计算与语言 · 计算机科学 2025-11-11 Mahmoud El-Bahnasawi

The shift of public debate to the digital sphere has been accompanied by a rise in online hate speech. While many promising approaches for hate speech classification have been proposed, studies often focus only on a single language, usually…

计算与语言 · 计算机科学 2023-01-03 Ana Kotarcic , Dominik Hangartner , Fabrizio Gilardi , Selina Kurer , Karsten Donnay

When building a predictive model, it is often difficult to ensure that application-specific requirements are encoded by the model that will eventually be deployed. Consider researchers working on hate speech detection. They will have an…

计算与语言 · 计算机科学 2025-01-14 Urja Khurana , Eric Nalisnick , Antske Fokkens

With the exponential rise in user-generated web content on social media, the proliferation of abusive languages towards an individual or a group across the different sections of the internet is also rapidly increasing. It is very…

计算与语言 · 计算机科学 2021-03-24 Prashant Kapil , Asif Ekbal

Social stereotypes negatively impact individuals' judgements about different groups and may have a critical role in how people understand language directed toward minority social groups. Here, we assess the role of social stereotypes in the…

计算与语言 · 计算机科学 2021-10-29 Aida Mostafazadeh Davani , Mohammad Atari , Brendan Kennedy , Morteza Dehghani

Existing work on automated hate speech detection typically focuses on binary classification or on differentiating among a small set of categories. In this paper, we propose a novel method on a fine-grained hate speech classification task,…

计算与语言 · 计算机科学 2018-09-05 Jing Qian , Mai ElSherief , Elizabeth Belding , William Yang Wang

Hate speech is a form of online harassment that involves the use of abusive language, and it is commonly seen in social media posts. This sort of harassment mainly focuses on specific group characteristics such as religion, gender,…

计算与语言 · 计算机科学 2022-06-10 Georgios K. Pitsilis

The rise of online platforms exacerbated the spread of hate speech, demanding scalable and effective detection. However, the accuracy of hate speech detection systems heavily relies on human-labeled data, which is inherently susceptible to…

计算与语言 · 计算机科学 2025-06-13 Tommaso Giorgi , Lorenzo Cima , Tiziano Fagni , Marco Avvenuti , Stefano Cresci

Our study addresses a significant gap in online hate speech detection research by focusing on homophobia, an area often neglected in sentiment analysis research. Utilising advanced sentiment analysis models, particularly BERT, and…

计算与语言 · 计算机科学 2024-05-16 Josh McGiff , Nikola S. Nikolov

Hate speech is a major issue in social networks due to the high volume of data generated daily. Recent works demonstrate the usefulness of machine learning (ML) in dealing with the nuances required to distinguish between hateful posts from…

计算与语言 · 计算机科学 2022-01-19 Rafael M. O. Cruz , Woshington V. de Sousa , George D. C. Cavalcanti

Online social media is rife with offensive and hateful comments, prompting the need for their automatic detection given the sheer amount of posts created every second. Creating high-quality human-labelled datasets for this task is difficult…

计算与语言 · 计算机科学 2023-08-01 João A. Leite , Carolina Scarton , Diego F. Silva

Estimating population parameters in finite populations of text documents can be challenging when obtaining the labels for the target variable requires manual annotation. To address this problem, we combine predictions from a transformer…

计算与语言 · 计算机科学 2025-05-09 Hannes Waldetoft , Jakob Torgander , Måns Magnusson

With the continuous growth of internet users and media content, it is very hard to track down hateful speech in audio and video. Converting video or audio into text does not detect hate speech accurately as human sometimes uses hateful…

人工智能 · 计算机科学 2023-07-24 Fariha Tahosin Boishakhi , Ponkoj Chandra Shill , Md. Golam Rabiul Alam

Current research on hate speech analysis is typically oriented towards monolingual and single classification tasks. In this paper, we present a new multilingual multi-aspect hate speech analysis dataset and use it to test the current…

计算与语言 · 计算机科学 2019-08-30 Nedjma Ousidhoum , Zizheng Lin , Hongming Zhang , Yangqiu Song , Dit-Yan Yeung

Hateful and Toxic content has become a significant concern in today's world due to an exponential rise in social media. The increase in hate speech and harmful content motivated researchers to dedicate substantial efforts to the challenging…

计算与语言 · 计算机科学 2021-01-25 Suman Dowlagar , Radhika Mamidi

The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a framework that explicitly…

计算与语言 · 计算机科学 2025-11-11 Brage Eilertsen , Røskva Bjørgfinsdóttir , Francielle Vargas , Ali Ramezani-Kebrya

As large language models achieve impressive scores on traditional benchmarks, an increasing number of researchers are becoming concerned about benchmark data leakage during pre-training, commonly known as the data contamination problem. To…

计算与语言 · 计算机科学 2024-06-27 Kun Qian , Shunji Wan , Claudia Tang , Youzhi Wang , Xuanming Zhang , Maximillian Chen , Zhou Yu