English
Related papers

Related papers: ToxiGen: A Large-Scale Machine-Generated Dataset f…

200 papers

User posts whose perceived toxicity depends on the conversational context are rare in current toxicity detection datasets. Hence, toxicity detectors trained on existing datasets will also tend to disregard context, making the detection of…

Computation and Language · Computer Science 2021-11-22 Alexandros Xenos , John Pavlopoulos , Ion Androutsopoulos , Lucas Dixon , Jeffrey Sorensen , Leo Laugier

Detoxification, the task of rewriting harmful language into non-toxic text, has become increasingly important amid the growing prevalence of toxic content online. However, high-quality parallel datasets for detoxification, especially for…

Computation and Language · Computer Science 2025-06-09 Shuzhou Yuan , Ercong Nie , Lukas Kouba , Ashish Yashwanth Kangen , Helmut Schmid , Hinrich Schütze , Michael Färber

Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomena such as toxicity, hate speech, or offensive language. In this work, we introduce…

Computation and Language · Computer Science 2026-05-12 Davide Bruni , Carlo Bardazzi , Maurizio Tesconi

Hate speech detection is key to online content moderation, but current models struggle to generalise beyond their training data. This has been linked to dataset biases and the use of sentence-level labels, which fail to teach models the…

Computation and Language · Computer Science 2025-06-05 Agostina Calabrese , Tom Sherborne , Björn Ross , Mirella Lapata

Attackers create adversarial text to deceive both human perception and the current AI systems to perform malicious purposes such as spam product reviews and fake political posts. We investigate the difference between the adversarial and the…

Computation and Language · Computer Science 2019-12-20 Hoang-Quoc Nguyen-Son , Tran Phuong Thao , Seira Hidano , Shinsaku Kiyomoto

Online hate on social media ranges from overt slurs and threats (\emph{hard hate speech}) to \emph{soft hate speech}: discourse that appears reasonable on the surface but uses framing and value-based arguments to steer audiences toward…

Computation and Language · Computer Science 2026-01-29 Xuanyu Su , Diana Inkpen , Nathalie Japkowicz

Hate speech detection is a common downstream application of natural language processing (NLP) in the real world. In spite of the increasing accuracy, current data-driven approaches could easily learn biases from the imbalanced data…

Computation and Language · Computer Science 2022-09-22 Yi Cai , Arthur Zimek , Gerhard Wunder , Eirini Ntoutsi

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication science theory, we…

Computation and Language · Computer Science 2026-03-25 Nils A. Herrmann , Tobias Eder , Jingyi He , Georg Groh

Detecting online hate is a complex task, and low-performing models have harmful consequences when used for sensitive applications such as content moderation. Emoji-based hate is an emerging challenge for automated detection. We present…

Computation and Language · Computer Science 2022-05-09 Hannah Rose Kirk , Bertram Vidgen , Paul Röttger , Tristan Thrush , Scott A. Hale

Detecting offensive language on social media is an important task. The ICWSM-2020 Data Challenge Task 2 is aimed at identifying offensive content using a crowd-sourced dataset containing 100k labelled tweets. The dataset, however, suffers…

Computation and Language · Computer Science 2020-12-08 Ruibo Liu , Guangxuan Xu , Soroush Vosoughi

In today's digital world, social media plays a significant role in facilitating communication and content sharing. However, the exponential rise in user-generated content has led to challenges in maintaining a respectful online environment.…

Computation and Language · Computer Science 2024-03-05 Mohammad Dehghani

One of the major challenges in automatic hate speech detection is the lack of datasets that cover a wide range of biased and unbiased messages and that are consistently labeled. We propose a labeling procedure that addresses some of the…

Computation and Language · Computer Science 2023-05-01 Gunther Jikeli , Sameer Karali , Daniel Miehling , Katharina Soemer

Automatic detection of hate and abusive language is essential to combat its online spread. Moreover, recognising and explaining hate speech serves to educate people about its negative effects. However, most current detection models operate…

Computation and Language · Computer Science 2025-05-06 Paloma Piot , Javier Parapar

The proliferation of hate speech has caused significant harm to society. The intensity and directionality of hate are closely tied to the target and argument it is associated with. However, research on hate speech detection in Chinese has…

Computation and Language · Computer Science 2025-05-21 Zewen Bai , Shengdi Yin , Junyu Lu , Jingjie Zeng , Haohao Zhu , Yuanyuan Sun , Liang Yang , Hongfei Lin

Implicit hate speech detection is challenging due to its subtlety and reliance on contextual interpretation rather than explicit offensive words. Current approaches rely on contrastive learning, which are shown to be effective on…

Computation and Language · Computer Science 2025-09-22 Yejin Lee , Joonghyuk Hahn , Hyeseon Ahn , Yo-Sub Han

Hate speech has grown significantly on social media, causing serious consequences for victims of all demographics. Despite much attention being paid to characterize and detect discriminatory speech, most work has focused on explicit or…

Computation and Language · Computer Science 2021-09-14 Mai ElSherief , Caleb Ziems , David Muchlinski , Vaishnavi Anupindi , Jordyn Seybolt , Munmun De Choudhury , Diyi Yang

As toxic language becomes nearly pervasive online, there has been increasing interest in leveraging the advancements in natural language processing (NLP), from very large transformer models to automatically detecting and removing toxic…

Computation and Language · Computer Science 2020-07-02 Austin P. Wright , Omar Shaikh , Haekyu Park , Will Epperson , Muhammed Ahmed , Stephane Pinel , Diyi Yang , Duen Horng Chau

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework…

Computation and Language · Computer Science 2023-05-24 Yiming Zhang , Sravani Nanduri , Liwei Jiang , Tongshuang Wu , Maarten Sap

Although various techniques have been proposed to generate adversarial samples for white-box attacks on text, little attention has been paid to black-box attacks, which are more realistic scenarios. In this paper, we present a novel…

Computation and Language · Computer Science 2018-05-24 Ji Gao , Jack Lanchantin , Mary Lou Soffa , Yanjun Qi

Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech and classify text…

Computation and Language · Computer Science 2022-03-25 Ke-Li Chiu , Annie Collins , Rohan Alexander