English
Related papers

Related papers: Toward Generalized Cross-Lingual Hateful Language …

200 papers

The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to generalize due to…

Computation and Language · Computer Science 2024-10-08 Huy Nghiem , Hal Daumé

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication science theory, we…

Computation and Language · Computer Science 2026-03-25 Nils A. Herrmann , Tobias Eder , Jingyi He , Georg Groh

Online harassment in the form of hate speech has been on the rise in recent years. Addressing the issue requires a combination of content moderation by people, aided by automatic detection methods. As content moderation is itself harmful to…

Computation and Language · Computer Science 2021-08-03 Sheikh Muhammad Sarwar , Vanessa Murdock

In support of open and reproducible research, there has been a rapidly increasing number of datasets made available for research. As the availability of datasets increases, it becomes more important to have quality metadata for discovering…

Computation and Language · Computer Science 2023-10-18 Shiwei Zhang , Mingfang Wu , Xiuzhen Zhang

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

Computation and Language · Computer Science 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring. Across 19 LLMs and two experiments totaling more than 4 million annotation judgments,…

Computation and Language · Computer Science 2026-03-17 Petter Törnberg

Although social media platforms are a prominent arena for users to engage in interpersonal discussions and express opinions, the facade and anonymity offered by social media may allow users to spew hate speech and offensive content. Given…

Computation and Language · Computer Science 2024-05-09 Ayushi Nirmal , Amrita Bhattacharjee , Paras Sheth , Huan Liu

Current disfluency detection methods heavily rely on costly and scarce human-annotated data. To tackle this issue, some approaches employ heuristic or statistical features to generate disfluent sentences, partially improving detection…

Computation and Language · Computer Science 2024-08-07 Zhenrong Cheng , Jiayan Guo , Hao Sun , Yan Zhang

Hate speech detection across contemporary social media presents unique challenges due to linguistic diversity and the informal nature of online discourse. These challenges are further amplified in settings involving code-mixing,…

Computation and Language · Computer Science 2025-06-17 Daman Deep Singh , Ramanuj Bhattacharjee , Abhijnan Chakraborty

Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and…

Computation and Language · Computer Science 2024-10-02 Lingzi Hong , Pengcheng Luo , Eduardo Blanco , Xiaoying Song

Social media is awash with hateful content, much of which is often veiled with linguistic and topical diversity. The benchmark datasets used for hate speech detection do not account for such divagation as they are predominantly compiled…

Computation and Language · Computer Science 2023-06-16 Atharva Kulkarni , Sarah Masud , Vikram Goyal , Tanmoy Chakraborty

Large Language Models (LLMs) offer a lucrative promise for scalable content moderation, including hate speech detection. However, they are also known to be brittle and biased against marginalised communities and dialects. This requires…

Computation and Language · Computer Science 2025-10-14 Ananya Malik , Kartik Sharma , Shaily Bhatt , Lynnette Hui Xian Ng

Detecting toxic content using language models is crucial yet challenging. While substantial progress has been made in English, toxicity detection in French remains underdeveloped, primarily due to the lack of culturally relevant,…

Computation and Language · Computer Science 2026-04-21 Axel Delaval , Shujian Yang , Haicheng Wang , Han Qiu , Jialiang Lu

Online marketplaces, while revolutionizing global commerce, have inadvertently facilitated the proliferation of illicit activities, including drug trafficking, counterfeit sales, and cybercrimes. Traditional content moderation methods such…

Computation and Language · Computer Science 2026-03-06 Quoc Khoa Tran , Thanh Thi Nguyen , Campbell Wilson

Multilingual large language models (LLMs) are known to more frequently generate non-faithful output in resource-constrained languages (Guerreiro et al., 2023 - arXiv:2303.16104), potentially because these typologically diverse languages are…

Computation and Language · Computer Science 2025-03-11 Tsan Tsai Chan , Xin Tong , Thi Thu Uyen Hoang , Barbare Tepnadze , Wojciech Stempniak

With the advance of large language models (LLMs), LLMs have been utilized for the various tasks. However, the issues of variability and reproducibility of results from each trial of LLMs have been largely overlooked in existing literature…

Computation and Language · Computer Science 2025-05-08 Junichiro Niimi

The massive spread of hate speech, hateful content targeted at specific subpopulations, is a problem of critical social importance. Automated methods of hate speech detection typically employ state-of-the-art deep learning (DL)-based text…

Computation and Language · Computer Science 2022-05-23 Tomer Wullach , Amir Adler , Einat Minkov

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

Computation and Language · Computer Science 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM…

The automatic identification of offensive language such as hate speech is important to keep discussions civil in online communities. Identifying hate speech in multimodal content is a particularly challenging task because offensiveness can…

Computation and Language · Computer Science 2024-02-20 Amrita Ganguly , Al Nahian Bin Emran , Sadiya Sayara Chowdhury Puspo , Md Nishat Raihan , Dhiman Goswami , Marcos Zampieri
‹ Prev 1 3 4 5 6 7 10 Next ›