English

Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation

Computation and Language 2025-06-25 v3

Abstract

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation, accuracy across diverse sensitive categories, and privacy concerns. Additionally, existing datasets and open-source models focus predominantly on toxic language, leaving gaps in detecting other sensitive categories such as substance abuse or self-harm. In this paper, we put forward a unified dataset tailored for social media content moderation across six sensitive categories: conflictual language, profanity, sexually explicit material, drug-related content, self-harm, and spam. By collecting and annotating data with consistent retrieval strategies and guidelines, we address the shortcomings of previous focalised research. Our analysis demonstrates that fine-tuning large language models (LLMs) on this novel dataset yields significant improvements in detection performance compared to open off-the-shelf models such as LLaMA, and even proprietary OpenAI models, which underperform by 10-15% overall. This limitation is even more pronounced on popular moderation APIs, which cannot be easily tailored to specific sensitive content categories, among others.

Keywords

Cite

@article{arxiv.2411.19832,
  title  = {Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation},
  author = {Dimosthenis Antypas and Indira Sen and Carla Perez-Almendros and Jose Camacho-Collados and Francesco Barbieri},
  journal= {arXiv preprint arXiv:2411.19832},
  year   = {2025}
}

Comments

Accepted at the 9th Workshop on Online Abuse and Harms (WOAH)

R2 v1 2026-06-28T20:17:02.407Z