English

Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection

Computation and Language 2024-10-22 v1

Abstract

Identifying offensive language is essential for maintaining safety and sustainability in the social media era. Though large language models (LLMs) have demonstrated encouraging potential in social media analytics, they lack thorough evaluation when in offensive language detection, particularly in multilingual environments. We for the first time evaluate multilingual offensive language detection of LLMs in three languages: English, Spanish, and German with three LLMs, GPT-3.5, Flan-T5, and Mistral, in both monolingual and multilingual settings. We further examine the impact of different prompt languages and augmented translation data for the task in non-English contexts. Furthermore, we discuss the impact of the inherent bias in LLMs and the datasets in the mispredictions related to sensitive topics.

Keywords

Cite

@article{arxiv.2410.15623,
  title  = {Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection},
  author = {Jianfei He and Lilin Wang and Jiaying Wang and Zhenyu Liu and Hongbin Na and Zimu Wang and Wei Wang and Qi Chen},
  journal= {arXiv preprint arXiv:2410.15623},
  year   = {2024}
}

Comments

Accepted at UIC 2024 proceedings. Accepted version