English
Related papers

Related papers: Digital Guardians: Can GPT-4, Perspective API, and…

200 papers

Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we…

Computation and Language · Computer Science 2025-11-06 Jay Patel , Hrudayangam Mehta , Jeremy Blackburn

This study evaluates the biases in Gemini 2.0 Flash Experimental, a state-of-the-art large language model (LLM) developed by Google, focusing on content moderation and gender disparities. By comparing its performance to ChatGPT-4o, examined…

Computation and Language · Computer Science 2025-03-24 Roberto Balestri

Automated hate speech detection is an important tool in combating the spread of hate speech, particularly in social media. Numerous methods have been developed for the task, including a recent proliferation of deep-learning based…

Computation and Language · Computer Science 2023-12-08 Jitendra Singh Malik , Hezhe Qiao , Guansong Pang , Anton van den Hengel

Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1…

Computation and Language · Computer Science 2021-09-09 Paul Röttger , Bertram Vidgen , Dong Nguyen , Zeerak Waseem , Helen Margetts , Janet B. Pierrehumbert

This study explores the robustness of university assessments against the use of Open AI's Generative Pre-Trained Transformer 4 (GPT-4) generated content and evaluates the ability of academic staff to detect its use when supported by the…

Computers and Society · Computer Science 2023-11-02 Mike Perkins , Jasper Roe , Darius Postma , James McGaughran , Don Hickerson

Combating online hate speech in multilingual settings requires approaches that go beyond English-centric models and capture the cultural and linguistic diversity of global online discourse. This paper presents a comprehensive survey and…

Computation and Language · Computer Science 2026-03-23 Zahra Safdari Fesaghandis , Suman Kalyan Maity

In the wake of a polarizing election, social media is laden with hateful content. To address various limitations of supervised hate speech classification methods including corpus bias and huge cost of annotation, we propose a weakly…

Computation and Language · Computer Science 2018-05-23 Lei Gao , Alexis Kuppersmith , Ruihong Huang

While recent research has focused on developing safeguards for generative AI (GAI) model-level content safety, little is known about how content moderation to prevent malicious content performs for end-users in real-world GAI products. To…

Human-Computer Interaction · Computer Science 2025-06-18 Lan Gao , Oscar Chen , Rachel Lee , Nick Feamster , Chenhao Tan , Marshini Chetty

The context-dependent nature of online aggression makes annotating large collections of data extremely difficult. Previously studied datasets in abusive language detection have been insufficient in size to efficiently train deep learning…

Computation and Language · Computer Science 2018-08-31 Younghun Lee , Seunghyun Yoon , Kyomin Jung

With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods…

Computation and Language · Computer Science 2021-01-19 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

We examined four case studies in the context of hate speech on Twitter in Italian from 2019 to 2020, aiming at comparing the classification of the 3,600 tweets made by expert pedagogists with the automatic classification made by machine…

Social and Information Networks · Computer Science 2024-02-14 Erica Forzinetti , Marco L. Della Vedova , Stefano Pasta , Milena Santerini

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often…

Multiagent Systems · Computer Science 2026-03-26 Rui Xing , Qi Chai , Jie Ma , Jing Tao , Pinghui Wang , Shuming Zhang , Xinping Wang , Hao Wang

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

Computation and Language · Computer Science 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

Computation and Language · Computer Science 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

The spread of information through social media platforms can create environments possibly hostile to vulnerable communities and silence certain groups in society. To mitigate such instances, several models have been developed to detect hate…

Computation and Language · Computer Science 2023-09-26 Marzieh Babaeianjelodar , Gurram Poorna Prudhvi , Stephen Lorenz , Keyu Chen , Sumona Mondal , Soumyabrata Dey , Navin Kumar

In the current era of the internet, where social media platforms are easily accessible for everyone, people often have to deal with threats, identity attacks, hate, and bullying due to their association with a cast, creed, gender, religion,…

Computation and Language · Computer Science 2021-12-21 Zaki Mustafa Farooqi , Sreyan Ghosh , Rajiv Ratn Shah

While the use of machine learning for the detection of propaganda techniques in text has garnered considerable attention, most approaches focus on "black-box" solutions with opaque inner workings. Interpretable approaches provide a…

Computation and Language · Computer Science 2024-07-17 Kyle Hamilton , Luca Longo , Bojan Bozic

Recently efforts have been made by social media platforms as well as researchers to detect hateful or toxic language using large language models. However, none of these works aim to use explanation, additional context and victim community…

Computation and Language · Computer Science 2023-10-31 Sarthak Roy , Ashish Harshavardhan , Animesh Mukherjee , Punyajoy Saha

The exponential increase in the use of the Internet and social media over the last two decades has changed human interaction. This has led to many positive outcomes, but at the same time it has brought risks and harms. While the volume of…

Computation and Language · Computer Science 2020-12-23 Neeraj Vashistha , Arkaitz Zubiaga , Shanky Sharma