English
Related papers

Related papers: A Quality Type-aware Annotated Corpus and Lexicon …

200 papers

Online hate speech is a recent problem in our society that is rising at a steady pace by leveraging the vulnerabilities of the corresponding regimes that characterise most social media platforms. This phenomenon is primarily fostered by…

Computation and Language · Computer Science 2022-01-05 Ioannis Mollas , Zoe Chrysopoulou , Stamatis Karlos , Grigorios Tsoumakas

This paper introduces a method for detecting inappropriately targeting language in online conversations by integrating crowd and expert annotations with ChatGPT. We focus on English conversation threads from Reddit, examining comments that…

Computation and Language · Computer Science 2025-05-23 Baran Barbarestani , Isa Maks , Piek Vossen

The detection of offensive, hateful and profane language has become a critical challenge since many users in social networks are exposed to cyberbullying activities on a daily basis. In this paper, we present an analysis of combining…

Computation and Language · Computer Science 2021-12-10 Sherzod Hakimov , Ralph Ewerth

Even though hate speech (HS) online has been an important object of research in the last decade, most HS-related corpora over-simplify the phenomenon of hate by attempting to label user comments as "hate" or "neutral". This ignores the…

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset…

Computation and Language · Computer Science 2020-05-29 Jae Yeon Kim , Carlos Ortiz , Sarah Nam , Sarah Santiago , Vivek Datta

Building a benchmark dataset for hate speech detection presents various challenges. Firstly, because hate speech is relatively rare, random sampling of tweets to annotate is very inefficient in finding hate speech. To address this, prior…

Computation and Language · Computer Science 2021-11-11 Md Mustafizur Rahman , Dinesh Balakrishnan , Dhiraj Murthy , Mucahid Kutlu , Matthew Lease

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

Crowdsourced annotation is vital to both collecting labelled data to train and test automated content moderation systems and to support human-in-the-loop review of system decisions. However, annotation tasks such as judging hate speech are…

Human-Computer Interaction · Computer Science 2023-09-06 Danula Hettiachchi , Indigo Holcombe-James , Stephanie Livingstone , Anjalee de Silva , Matthew Lease , Flora D. Salim , Mark Sanderson

Hate speech is a form of online harassment that involves the use of abusive language, and it is commonly seen in social media posts. This sort of harassment mainly focuses on specific group characteristics such as religion, gender,…

Computation and Language · Computer Science 2022-06-10 Georgios K. Pitsilis

The ability to accurately detect and filter offensive content automatically is important to ensure a rich and diverse digital discourse. Trolling is a type of hurtful or offensive content that is prevalent in social media, but is…

Computers and Society · Computer Science 2020-08-04 Hitkul , Karmanya Aggarwal , Pakhi Bamdev , Debanjan Mahata , Rajiv Ratn Shah , Ponnurangam Kumaraguru

With rising concern around abusive and hateful behavior on social media platforms, we present an ensemble learning method to identify and analyze the linguistic properties of such content. Our stacked ensemble comprises of three machine…

Computation and Language · Computer Science 2020-06-08 Gaurav Verma , Niyati Chhaya , Vishwa Vinay

Harmful speech has various forms and it has been plaguing the social media in different ways. If we need to crackdown different degrees of hate speech and abusive behavior amongst it, the classification needs to be based on complex…

Computation and Language · Computer Science 2018-06-13 Sanjana Sharma , Saksham Agrawal , Manish Shrivastava

The datasets most widely used for abusive language detection contain lists of messages, usually tweets, that have been manually judged as abusive or not by one or more annotators, with the annotation performed at message level. In this…

Computation and Language · Computer Science 2021-03-30 Stefano Menini , Alessio Palmero Aprosio , Sara Tonelli

Hate speech is a specific type of controversial content that is widely legislated as a crime that must be identified and blocked. However, due to the sheer volume and velocity of the Twitter data stream, hate speech detection cannot be…

Computation and Language · Computer Science 2021-08-09 Moin Khan , Khurram Shahzad , Kamran Malik

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and abusive language.…

Computation and Language · Computer Science 2019-05-30 Thomas Davidson , Debasmita Bhattacharya , Ingmar Weber

Identifying misogyny using artificial intelligence is a form of combating online toxicity against women. However, the subjective nature of interpreting misogyny poses a significant challenge to model the phenomenon. In this paper, we…

Computation and Language · Computer Science 2024-06-25 Jason Angel , Segun Taofeek Aroyehun , Grigori Sidorov , Alexander Gelbukh

This paper presents a novel scheme for the annotation of hate speech in corpora of Web 2.0 commentary. The proposed scheme is motivated by the critical analysis of posts made in reaction to news reports on the Mediterranean migration crisis…

Computers and Society · Computer Science 2020-08-17 Stavros Assimakopoulos , Rebecca Vella Muskat , Lonneke van der Plas , Albert Gatt

Online toxic content has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. A significant amount of research has been focused on detecting or analyzing toxic content using machine-learning…

Computation and Language · Computer Science 2025-09-19 Gautam Kishore Shahi , Tim A. Majchrzak

Hate speech, offensive language, sexism, racism and other types of abusive behavior have become a common phenomenon in many online social media platforms. In recent years, such diverse abusive behaviors have been manifesting with increased…

Computation and Language · Computer Science 2018-02-22 Antigoni-Maria Founta , Despoina Chatzakou , Nicolas Kourtellis , Jeremy Blackburn , Athena Vakali , Ilias Leontiadis

In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on…

Computation and Language · Computer Science 2020-03-18 Shiladitya Bhattacharya , Siddharth Singh , Ritesh Kumar , Akanksha Bansal , Akash Bhagat , Yogesh Dawer , Bornini Lahiri , Atul Kr. Ojha