English

SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection

Artificial Intelligence 2026-03-20 v3 Computation and Language Computers and Society

Abstract

We introduce SynBullying, a synthetic multi-LLM conversational dataset for studying and detecting cyberbullying (CB). SynBullying provides a scalable and ethically safe alternative to human data collection by leveraging large language models (LLMs) to simulate realistic bullying interactions. The dataset offers (i) conversational structure, capturing multi-turn exchanges rather than isolated posts; (ii) context-aware annotations, where harmfulness is assessed within the conversational flow considering context, intent, and discourse dynamics; and (iii) fine-grained labeling, covering various CB categories for detailed linguistic and behavioral analysis. We evaluate SynBullying across five dimensions, including conversational structure, lexical patterns, sentiment/toxicity, role dynamics, harm intensity, and CB-type distribution. We further examine its utility by testing its performance as standalone training data and as an augmentation source for CB classification.

Keywords

Cite

@article{arxiv.2511.11599,
  title  = {SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detection},
  author = {Arefeh Kazemi and Hamza Qadeer and Joachim Wagner and Hossein Hosseini and Sri Balaaji Natarajan Kalaivendan and Brian Davis},
  journal= {arXiv preprint arXiv:2511.11599},
  year   = {2026}
}