English
Related papers

Related papers: Harmful Speech Detection by Language Models Exhibi…

200 papers

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence, graphic content,…

Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses…

Computers and Society · Computer Science 2025-06-06 Emma Harvey , Rene F. Kizilcec , Allison Koenecke

Biased news contributes to societal polarization and is often reinforced by hostile reader comments, constituting a vital yet often overlooked aspect of news dissemination. Our study reveals that offensive comments support biased content,…

Computation and Language · Computer Science 2025-08-25 Luyang Lin , Zijin Feng , Lingzhi Wang , Kam-Fai Wong

Many social science questions ask how linguistic properties causally affect an audience's attitudes and behaviors. Because text properties are often interlinked (e.g., angry reviews use profane language), we must control for possible latent…

Recent advances in large language models (LLMs) have demonstrated strong performance on simple text classification tasks, frequently under zero-shot settings. However, their efficacy declines when tackling complex social media challenges…

Computation and Language · Computer Science 2025-04-23 Elyas Meguellati , Assaad Zeghina , Shazia Sadiq , Gianluca Demartini

Abusive language detection has become an increasingly important task as a means to tackle this type of harmful content in social media. There has been a substantial body of research developing models for determining if a social media post…

Computation and Language · Computer Science 2025-08-19 Raneem Alharthi , Rajwa Alharthi , Aiqi Jiang , Arkaitz Zubiaga

Large language models (LLMs) excel in many diverse applications beyond language generation, e.g., translation, summarization, and sentiment analysis. One intriguing application is in text classification. This becomes pertinent in the realm…

Computation and Language · Computer Science 2024-03-14 Tharindu Kumarage , Amrita Bhattacharjee , Joshua Garland

Hate speech in social media is a growing phenomenon, and detecting such toxic content has recently gained significant traction in the research community. Existing studies have explored fine-tuning language models (LMs) to perform hate…

Computation and Language · Computer Science 2023-03-07 Md Rabiul Awal , Roy Ka-Wei Lee , Eshaan Tanwar , Tanmay Garg , Tanmoy Chakraborty

In the evolving landscape of online communication, moderating hate speech (HS) presents an intricate challenge, compounded by the multimodal nature of digital content. This comprehensive survey delves into the recent strides in HS…

Computation and Language · Computer Science 2024-10-31 Ming Shan Hee , Shivam Sharma , Rui Cao , Palash Nandi , Preslav Nakov , Tanmoy Chakraborty , Roy Ka-Wei Lee

Online memes are a powerful yet challenging medium for content moderation, often masking harmful intent behind humor, irony, or cultural symbolism. Conventional moderation systems "especially those relying on explicit text" frequently fail…

Information Retrieval · Computer Science 2025-10-20 Sayantan Adak , Somnath Banerjee , Rajarshi Mandal , Avik Halder , Sayan Layek , Rima Hazra , Animesh Mukherjee

Social media continues to have an impact on the trajectory of humanity. However, its introduction has also weaponized keyboards, allowing the abusive language normally reserved for in-person bullying to jump onto the screen, i.e.,…

Machine Learning · Computer Science 2024-09-20 Muhammad Arslan , Manuel Sandoval Madrigal , Mohammed Abuhamad , Deborah L. Hall , Yasin N. Silva

Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment. This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content…

Computation and Language · Computer Science 2024-06-19 Yunze Xiao , Yujia Hu , Kenny Tsu Wei Choo , Roy Ka-wei Lee

Despite the valuable social interactions that online media promote, these systems provide space for speech that would be potentially detrimental to different groups of people. The moderation of content imposed by many social media has…

Social and Information Networks · Computer Science 2021-08-30 Lucas Henrique Costa de Lima , Julio Reis , Philipe Melo , Fabricio Murai , Fabricio Benevenuto

In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election…

Online conversations can be toxic and subjected to threats, abuse, or harassment. To identify toxic text comments, several deep learning and machine learning models have been proposed throughout the years. However, recent studies…

Machine Learning · Computer Science 2023-11-09 Md Azim Khan

We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community. The benchmark is community-sourced, via application of a novel method that…

Computation and Language · Computer Science 2024-10-21 Virginia K. Felkner , Ho-Chun Herbert Chang , Eugene Jang , Jonathan May

The rapid deployment of generative language models (LMs) has raised concerns about social biases affecting the well-being of diverse consumers. The extant literature on generative LMs has primarily examined bias via explicit identity…

Computation and Language · Computer Science 2026-05-04 Evan Shieh , Faye-Marie Vassel , Cassidy Sugimoto , Thema Monroe-White

Large Language Models (LLMs) are increasingly deployed to curate and rank human-created content, yet the nature and structure of their biases in these tasks remains poorly understood: which biases are robust across providers and platforms,…

Social and Information Networks · Computer Science 2026-04-20 Nicolò Pagan , Christopher Barrie , Chris Andrew Bail , Petter Törnberg

Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic…

Computation and Language · Computer Science 2025-05-22 Xintong Wang , Yixiao Liu , Jingheng Pan , Liang Ding , Longyue Wang , Chris Biemann

With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency,…

Computation and Language · Computer Science 2024-09-27 Guanyi Mou , Kyumin Lee