English
Related papers

Related papers: How We Define Harm Impacts Data Annotations: Expla…

200 papers

Hate Speech takes many forms to target communities with derogatory comments, and takes humanity a step back in societal progress. HateXplain is a recently published and first dataset to use annotated spans in the form of rationales, along…

Computation and Language · Computer Science 2022-08-10 Arvind Subramaniam , Aryan Mehra , Sayani Kundu

Toxic comment classification models are often found biased toward identity terms which are terms characterizing a specific group of people such as "Muslim" and "black". Such bias is commonly reflected in false-positive predictions, i.e.…

Computation and Language · Computer Science 2022-10-18 Zhixue Zhao , Ziqi Zhang , Frank Hopfgartner

Social media platforms are plagued by harmful content such as hate speech, misinformation, and extremist rhetoric. Machine learning (ML) models are widely adopted to detect such content; however, they remain highly vulnerable to adversarial…

Machine Learning · Computer Science 2025-12-30 Yidong Chai , Yi Liu , Mohammadreza Ebrahimi , Weifeng Li , Balaji Padmanabhan

Despite the extensive communication benefits offered by social media platforms, numerous challenges must be addressed to ensure user safety. One of the most significant risks faced by users on these platforms is targeted hate speech. Social…

Computation and Language · Computer Science 2024-07-18 Sadar Jaf , Basel Barakat

Hate speech classifiers trained on imbalanced datasets struggle to determine if group identifiers like "gay" or "black" are used in offensive or prejudiced ways. Such biases manifest in false positives when these identifiers are present,…

Computation and Language · Computer Science 2020-07-08 Brendan Kennedy , Xisen Jin , Aida Mostafazadeh Davani , Morteza Dehghani , Xiang Ren

As autonomous systems rapidly become ubiquitous, there is a growing need for a legal and regulatory framework to address when and how such a system harms someone. There have been several attempts within the philosophy literature to define…

Artificial Intelligence · Computer Science 2023-01-20 Sander Beckers , Hana Chockler , Joseph Y. Halpern

When reading news articles on social networking services and news sites, readers can view comments marked by other people on these articles. By reading these comments, a reader can understand the public opinion about the news, and it is…

Computation and Language · Computer Science 2022-12-27 Teruki Nakahara , Taketoshi Ushiama

With the increasing diversity of use cases of large language models, a more informative treatment of texts seems necessary. An argumentative analysis could foster a more reasoned usage of chatbots, text completion mechanisms or other…

Computation and Language · Computer Science 2023-06-06 Damián Furman , Pablo Torres , José A. Rodríguez , Diego Letzen , Vanina Martínez , Laura Alonso Alemany

The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise. For toxic text identification, task-specific fine-tuning of these models are performed using datasets labeled by annotators who provide…

Computation and Language · Computer Science 2021-12-08 Kofi Arhin , Ioana Baldini , Dennis Wei , Karthikeyan Natesan Ramamurthy , Moninder Singh

Socio-linguistic indicators of affectively-relevant phenomena, such as emotion or sentiment, are often extracted from text to better understand features of human-computer interactions, including on social media. However, an indicator that…

Machine Learning · Computer Science 2025-11-24 Keith Burghardt , Daniel M. T. Fessler , Chyna Tang , Anne Pisor , Kristina Lerman

Current annotation agreement metrics are not well-suited for inter-group analysis, are sensitive to group size imbalances and restricted to single-annotation settings. These restrictions render them insufficient for many subjective tasks…

Computation and Language · Computer Science 2026-02-09 Dimitris Tsirmpas , John Pavlopoulos

NeurIPS 2020 requested that research paper submissions include impact statements on "potential nefarious uses and the consequences of failure." However, as researchers, practitioners and system designers, a key challenge to anticipating…

Computers and Society · Computer Science 2020-12-11 Margarita Boyarskaya , Alexandra Olteanu , Kate Crawford

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset…

Computation and Language · Computer Science 2020-05-29 Jae Yeon Kim , Carlos Ortiz , Sarah Nam , Sarah Santiago , Vivek Datta

We examined four case studies in the context of hate speech on Twitter in Italian from 2019 to 2020, aiming at comparing the classification of the 3,600 tweets made by expert pedagogists with the automatic classification made by machine…

Social and Information Networks · Computer Science 2024-02-14 Erica Forzinetti , Marco L. Della Vedova , Stefano Pasta , Milena Santerini

Social media algorithms are thought to amplify variation in user beliefs, thus contributing to radicalization. However, quantitative evidence on how algorithms and user preferences jointly shape harmful online engagement is limited. I…

General Economics · Economics 2025-03-11 Aarushi Kalra

This paper presents the contribution of the Data Science Kitchen at GermEval 2021 shared task on the identification of toxic, engaging, and fact-claiming comments. The task aims at extending the identification of offensive language, by…

Computation and Language · Computer Science 2024-08-20 Niclas Hildebrandt , Benedikt Boenninghoff , Dennis Orth , Christopher Schymura

Hate speech remains a persistent and unresolved challenge in online platforms. Content moderators, working on the front lines to review user-generated content and shield viewers from hate speech, often find themselves unprotected from the…

Human-Computer Interaction · Computer Science 2025-08-04 Subin Park , Jeonghyun Kim , Jeanne Choi , Joseph Seering , Uichin Lee , Sung-Ju Lee

Online harms are a growing problem in digital spaces, putting user safety at risk and reducing trust in social media platforms. One of the most persistent forms of harm is hate speech. To address this, we need tools that combine the speed…

Computation and Language · Computer Science 2025-09-03 Paloma Piot , Diego Sánchez , Javier Parapar

Warning: This paper contains examples of the language that some people may find offensive. Detecting and reducing hateful, abusive, offensive comments is a critical and challenging task on social media. Moreover, few studies aim to mitigate…

Computation and Language · Computer Science 2023-12-21 Neeraj Kumar Singh , Koyel Ghosh , Joy Mahapatra , Utpal Garain , Apurbalal Senapati

With surge in online platforms, there has been an upsurge in the user engagement on these platforms via comments and reactions. A large portion of such textual comments are abusive, rude and offensive to the audience. With machine learning…

Computation and Language · Computer Science 2021-08-17 Ayush Kumar , Pratik Kumar